Safety & Alignment
aka RL from AI feedback
RLAIF
Replacing human preference labels in RLHF with judgments from another model.
Definition
RLAIF uses a strong reference model to rank candidate outputs instead of human labellers, then trains a reward model and runs PPO as in RLHF. It scales cheaply and, with careful prompts, matches RLHF quality on alignment tasks. Constitutional AI is a notable RLAIF variant.
Common use cases
- Cheap alignment
- Scaling preference data
- Self-improvement