Safety & Alignment
Harmful Content Filter
Classifier that blocks or modifies model outputs containing disallowed content.
Definition
Harmful-content filters â Moderation API, Perspective API, Azure Content Safety â sit between the model and the user, scoring outputs for categories like hate, violence, sexual content and self-harm. They are a second line of defence beyond refusal training.
Common use cases
- Trust & safety
- Compliance
- Content moderation