Safety & Alignment

Harmful Content Filter

Classifier that blocks or modifies model outputs containing disallowed content.

Definition

Harmful-content filters — Moderation API, Perspective API, Azure Content Safety — sit between the model and the user, scoring outputs for categories like hate, violence, sexual content and self-harm. They are a second line of defence beyond refusal training.

Common use cases

  • Trust & safety
  • Compliance
  • Content moderation

Related terms

    Harmful Content Filter — AI Glossary | Railwail