Technique
Batch Inference
Processing many requests together â async, off-peak â at sharply reduced per-token cost.
Definition
Batch inference APIs (OpenAI, Anthropic, Google) queue submitted requests and process them within 24h at 30â50% lower price. They suit non-interactive workloads like ETL, evals or back-fills where latency does not matter.
Common use cases
- ETL
- Evals
- Bulk labelling