Safety & Alignment

Context Window Attack

Adversarial payload designed to overflow or exploit a model's context window.

Definition

Long context attacks exploit drift, repeated jailbreak phrases, or hidden instructions buried deep in retrieved documents. Defences include input length limits, recency-weighted attention monitoring and adversarial training.

Common use cases

  • Threat modelling
  • RAG security
  • Red-teaming

Related terms

    Context Window Attack — AI Glossary | Railwail