Technique

Batch Inference

Processing many requests together — async, off-peak — at sharply reduced per-token cost.

Definition

Batch inference APIs (OpenAI, Anthropic, Google) queue submitted requests and process them within 24h at 30–50% lower price. They suit non-interactive workloads like ETL, evals or back-fills where latency does not matter.

Common use cases

  • ETL
  • Evals
  • Bulk labelling

Related terms

    Batch Inference — AI Glossary | Railwail