Safety & Alignment
Harmful Content Filter
Classifier that blocks or modifies model outputs containing disallowed content.
Definition
Harmful-content filters — Moderation API, Perspective API, Azure Content Safety — sit between the model and the user, scoring outputs for categories like hate, violence, sexual content and self-harm. They are a second line of defence beyond refusal training.
Common use cases
- Trust & safety
- Compliance
- Content moderation