Safety & Alignment

Alignment

Engineering models so their behaviour matches human values, intent, and safety policies.

Definition

Alignment is the discipline of making AI systems pursue intended goals reliably and safely. Techniques span supervised fine-tuning, RLHF, RLAIF, constitutional AI, red-teaming and interpretability. It is the central concern of frontier labs and AI safety organisations.

Common use cases

  • Chat safety
  • Policy enforcement
  • Bias reduction

Related terms

    Alignment — AI Glossary | Railwail