Safety & Alignment
aka RL from AI feedback

RLAIF

Replacing human preference labels in RLHF with judgments from another model.

Definition

RLAIF uses a strong reference model to rank candidate outputs instead of human labellers, then trains a reward model and runs PPO as in RLHF. It scales cheaply and, with careful prompts, matches RLHF quality on alignment tasks. Constitutional AI is a notable RLAIF variant.

Common use cases

  • Cheap alignment
  • Scaling preference data
  • Self-improvement

Related terms

    RLAIF — AI Glossary | Railwail