LLM API Cost Calculator
Compare the monthly cost of 143+ LLM APIs from a single screen. Drag the sliders, see your exact bill.
Pricing AI workloads is harder than it should be. Every provider quotes a different unit β per 1K tokens, per 1M tokens, per call, per second of audio β and most published rate cards leave out fixed costs, batching discounts, and image surcharges. This calculator normalises 409+ models from Railwail's catalogue so you can compare them apples-to-apples in seconds.
How the math works
Each model has a per-1K-token input price and a per-1K-token output price, stored in credits (1 credit = $0.01). Your monthly bill is simply input_tokens Γ input_price + output_tokens Γ output_price. We surface the per-1M numbers in the table because most providers quote that unit publicly, but the underlying math runs on per-1K. Fixed per-call costs (image generation, video, audio) don't apply to token-based models and are excluded from this view β see /pricing for per-call workloads.
Switch providers, keep the same code
All models in the table speak the same OpenAI-compatible API on Railwail. Once you've identified a cheaper option, swapping is one environment-variable change β no SDK rewrite, no retry logic. Compare any two models head-to-head on the /compare page, or browse the full catalogue at /models.
Switching from GPT-5.5 to OpenAI text-embedding-3-small would save you ~100% on a 1M-input + 1M-output workload β same monthly traffic, same prompts, just a different model behind the API.
The prompt + context you send to the model each month.
The total response volume the model generates.
Token-based categories only β image, video, audio use per-call pricing.
Filter the table to one upstream provider.
| Action | |||||
|---|---|---|---|---|---|
BGE Large EN v1.5 huggingface | β | β | n/a | 512 | Try |
BGE-M3 (Multilingual) huggingface | β | β | n/a | 8K | Try |
Bio_ClinicalBERT huggingface | β | β | n/a | β | Try |
Biomedical NER (all entities) huggingface | β | β | n/a | β | Try |
| β | β | n/a | 200K | Try | |
| β | β | n/a | 200K | Try | |
| β | β | n/a | 200K | Try | |
| β | β | n/a | 256K | Try | |
| β | β | n/a | 131K | Try | |
ESM-2 650M (Protein Embeddings) huggingface | β | β | n/a | 1K | Try |
| β | β | n/a | 2.1M | Try | |
| β | β | n/a | 1.0M | Try | |
| β | β | n/a | 1.0M | Try | |
| β | β | n/a | 1.0M | Try | |
| β | β | n/a | 131K | Try | |
Medical NER (DeBERTa) huggingface | β | β | n/a | β | Try |
| β | β | n/a | 4.1M | Try | |
Nomic Embed Text v1.5 huggingface | β | β | n/a | 8K | Try |
| β | β | n/a | 200K | Try | |
PubMedBERT Embeddings (NeuML) huggingface | β | β | n/a | 512 | Try |
SPECTER (Scientific Paper Embeddings) huggingface | β | β | n/a | 512 | Try |
| β | β | n/a | 32K | Try | |
| β | β | n/a | 256K | Try | |
| β | β | n/a | 256K | Try | |
BLIP Image Captioning Large huggingface | β | β | n/a | β | Try |
BioBERT Disease NER (NCBI) huggingface | β | β | n/a | β | Try |
BioBERT v1.2 (Biomedical Embeddings) huggingface | β | β | n/a | 512 | Try |
BiomedBERT (PubMedBERT abstract) huggingface | β | β | n/a | 512 | Try |
| β | β | n/a | 200K | Try | |
Clinical Assertion and Negation BERT huggingface | β | β | n/a | β | Try |
Clinical NER (problem, test, treatment) huggingface | β | β | n/a | β | Try |
CodeGen 350M Mono huggingface | β | β | n/a | 2K | Try |
| β | β | n/a | 8K | Try | |
| β | β | n/a | 4K | Try | |
| β | β | n/a | 128K | Try | |
| β | β | n/a | 128K | Try | |
| β | β | n/a | 512 | Try | |
DeepSeek Coder 1.3B Instruct huggingface | β | β | n/a | 16K | Try |
| β | β | n/a | 128K | Try | |
| β | β | n/a | 64K | Try | |
| β | β | n/a | 64K | Try | |
| β | β | n/a | 1.1M | Try | |
| β | β | n/a | 1.1M | Try | |
| β | β | n/a | 1.1M | Try | |
GTE Large EN v1.5 huggingface | β | β | n/a | 8K | Try |
| β | β | n/a | 1.0M | Try | |
| β | β | n/a | 1.0M | Try | |
| β | β | n/a | 1.0M | Try | |
| β | β | n/a | 256K | Try | |
| β | β | n/a | 8K | Try | |
| β | β | n/a | 131K | Try | |
| β | β | n/a | β | Try | |
| β | β | n/a | 128K | Try | |
Multilingual E5 Large huggingface | β | β | n/a | 512 | Try |
| β | β | n/a | 127K | Try | |
| β | β | n/a | 127K | Try | |
| β | β | n/a | 131K | Try | |
| β | β | n/a | 33K | Try | |
Qwen2.5-Coder 32B Instruct huggingface | β | β | n/a | 131K | Try |
Qwen2.5-Coder 7B Instruct huggingface | β | β | n/a | 131K | Try |
Qwen2.5-VL 7B Instruct (HF) huggingface | β | β | n/a | 33K | Try |
| β | β | n/a | 128K | Try | |
| β | β | n/a | 16K | Try | |
| β | β | n/a | 128K | Try | |
Replit Code v1.5 3B huggingface | β | β | n/a | 4K | Try |
SciBERT (scivocab uncased) huggingface | β | β | n/a | 512 | Try |
| β | β | n/a | 4K | Try | |
Stable Code Instruct 3B huggingface | β | β | n/a | 16K | Try |
| β | β | n/a | 32K | Try | |
| β | β | n/a | 33K | Try | |
| β | β | n/a | 512 | Try | |
| β | β | $0.0003 | β | Try | |
MiDaS v3.1 Replicate Cheapest | β | β | $0.0003 | β | Try |
Segformer B5 Replicate Cheapest | β | β | $0.0003 | β | Try |
| β | β | $0.0006 | β | Try | |
| β | β | $0.0012 | β | Try | |
| β | β | $0.0012 | β | Try | |
| β | β | $0.0012 | β | Try | |
| β | β | $0.0012 | 8K | Try | |
| β | β | $0.0012 | β | Try | |
| β | β | $0.0012 | 4K | Try | |
| β | β | $0.0014 | β | Try | |
| β | β | $0.0015 | β | Try | |
| β | β | $0.0015 | 33K | Try | |
| β | β | $0.0020 | β | Try | |
| $0.024 | β | $0.0024 | 8K | Try | |
| β | β | $0.0024 | β | Try | |
| β | β | $0.0030 | 16K | Try | |
| β | β | $0.0039 | β | Try | |
| β | β | $0.0050 | β | Try | |
| β | β | $0.0065 | 16K | Try | |
| β | β | $0.0070 | 131K | Try | |
| β | β | $0.0074 | 16K | Try | |
| β | β | $0.0086 | 4K | Try | |
| β | β | $0.0094 | β | Try | |
| β | β | $0.011 | 8K | Try | |
| β | β | $0.013 | β | Try | |
| $0.156 | β | $0.016 | 8K | Try | |
| β | β | $0.018 | β | Try | |
| β | β | $0.020 | β | Try | |
| β | β | $0.022 | β | Try | |
| β | β | $0.025 | 16K | Try | |
| β | β | $0.028 | β | Try | |
| $0.060 | $0.300 | $0.036 | 128K | Try | |
| β | β | $0.041 | 16K | Try | |
| β | β | $0.041 | β | Try | |
| β | β | $0.046 | β | Try | |
| β | β | $0.050 | 16K | Try | |
| β | β | $0.050 | β | Try | |
| β | β | $0.065 | β | Try | |
| $0.120 | $0.600 | $0.072 | 8K | Try | |
| $0.180 | $0.720 | $0.090 | 128K | Try | |
| $0.180 | $0.720 | $0.090 | 128K | Try | |
| β | β | $0.092 | 16K | Try | |
| β | β | $0.094 | β | Try | |
| β | β | $0.102 | 4K | Try | |
| $0.240 | $1.50 | $0.174 | 400K | Try | |
| $0.360 | $1.44 | $0.180 | 1.0M | Try | |
| $0.360 | $1.44 | $0.180 | 1.0M | Try | |
| β | β | $0.216 | 16K | Try | |
| $0.300 | $2.40 | $0.270 | 400K | Try | |
| $0.600 | $3.60 | $0.420 | 1.0M | Try | |
| $0.900 | $5.40 | $0.630 | 400K | Try | |
| $1.58 | $4.75 | $0.634 | 1.0M | Try | |
| $1.32 | $5.28 | $0.660 | 200K | Try | |
| $1.32 | $5.28 | $0.660 | 200K | Try | |
| $1.20 | $6.00 | $0.720 | 200K | Try | |
| β | β | $1.19 | 16K | Try | |
| $2.40 | $9.60 | $1.20 | 1.0M | Try | |
| $2.40 | $9.60 | $1.20 | 200K | Try | |
| $1.50 | $12.00 | $1.35 | 1.0M | Try | |
| $1.50 | $12.00 | $1.35 | 400K | Try | |
| $2.40 | $12.00 | $1.44 | 1.0M | Try | |
| $3.00 | $12.00 | $1.50 | 128K | Try | |
| $3.00 | $12.00 | $1.50 | 128K | Try | |
| $2.40 | $14.40 | $1.68 | 1.0M | Try | |
| $3.00 | $18.00 | $2.10 | 1.1M | Try | |
| $3.60 | $18.00 | $2.16 | 1.0M | Try | |
| $4.80 | $24.00 | $2.88 | 1.0M | Try | |
| $6.00 | $30.00 | $3.60 | 1.0M | Try | |
| $6.00 | $30.00 | $3.60 | 1.0M | Try | |
| $6.00 | $36.00 | $4.20 | 400K | Try | |
| $12.00 | $60.00 | $7.20 | 1.0M | Try |
Cheapest 5 models today
Ranked by combined input + output cost for a 1M + 1M token monthly workload.
Frequently asked questions
How is the monthly cost calculated?
For each model we apply (input_tokens Γ input_price + output_tokens Γ output_price) / 1000. Prices are stored in credits per 1K tokens (1 credit = $0.01); the table shows the USD per-1M equivalent because providers quote that unit publicly.
Why is OpenAI more expensive than open-weights alternatives?
Closed-frontier models bake R&D, alignment work, and hosted inference into a single price. Open-weights models (Llama, Mistral, DeepSeek) push the inference cost to hosting providers competing on margin, which is why per-token prices can be 5-50Γ lower for comparable quality.
What's the difference between input and output cost?
Input tokens are the prompt you send; output tokens are the model's response. Output is typically 2-5Γ more expensive because generation is sequential and dominates GPU time. Reasoning models (o1, DeepSeek R1) charge for hidden "thinking" tokens as output, which is the main cost driver.
Does this include batch discounts or volume pricing?
No β list price only. Most providers offer 25-50% discounts for batch endpoints and enterprise contracts; use the table as a ceiling, not a floor.
Sign up to start using these models
One OpenAI-compatible endpoint, every model in the table above. Free credits to start, transparent per-token pricing thereafter.