Safety & Alignment

Refusal

Model behaviour of declining a request that violates its safety policy.

Definition

Refusal training teaches a model to recognise unsafe requests and respond with explanations or alternatives instead. Tuning the threshold is delicate: over-refusal harms helpfulness, under-refusal exposes risk. Refusal benchmarks (XSTest, MaliciousInstruct) probe this balance.

Common use cases

  • Safety filters
  • Compliance
  • Risk assessment

Related terms

    Refusal — AI Glossary | Railwail