Best AI language models, ranked by quality
Language models for chat, writing, coding and reasoning, compared on the quality people vote for in the Arena, the price of 1M tokens and quality per dollar.
Best AI language models, ranked by quality
Sorted by Arena score, highest first. Models without an Arena rating follow, by price.
Median Arena score here: 1,442
| # | Model | Arena score | Price | Value | Actions |
|---|---|---|---|---|---|
| 01 | Arena1,494±4 · 61,128 votes | US$12.00 / 1M tokensUS$6.00 in · US$30.00 out | 4 pts/$+52 over median | Try | |
| 02 | Gemini 3.1 Pro Google DeepMind | Arena1,487±3 · 106,951 votes | US$5.40 / 1M tokensUS$2.40 in · US$14.40 out | 8 pts/$+45 over median | Try |
| 03 | GPT-5.5 OpenAI | Arena1,476±4 · 66,317 votes | US$13.50 / 1M tokensUS$6.00 in · US$36.00 out | 3 pts/$+34 over median | Try |
| 04 | Gemini 3 Flash Google DeepMindLegacy | Arena1,474±4 · 30,225 votes | US$1.35 / 1M tokensUS$0.60 in · US$3.60 out | 23 pts/$+31 over median | Try |
| 05 | Arena1,473±4 · 53,446 votes | US$12.00 / 1M tokensUS$6.00 in · US$30.00 out | 3 pts/$+31 over median | Try | |
| 06 | Arena1,473±4 · 66,208 votes | US$7.20 / 1M tokensUS$3.60 in · US$18.00 out | 4 pts/$+30 over median | Try | |
| 07 | GPT-5.4 OpenAI | Arena1,466±4 · 63,526 votes | US$6.75 / 1M tokensUS$3.00 in · US$18.00 out | 3 pts/$+23 over median | Try |
| 08 | Gemini 2.5 Pro Google DeepMindLegacy | Arena1,446±3 · 122,554 votes | US$4.125 / 1M tokensUS$1.50 in · US$12.00 out | 1 pts/$+3 over median | Try |
| 09 | GPT-5.1 OpenAI | Arena1,439±4 · 42,982 votes | US$4.125 / 1M tokensUS$1.50 in · US$12.00 out | Not clearly above the median | Try |
| 10 | OpenAI o3 OpenAILegacy | Arena1,432±4 · 58,579 votes | US$4.20 / 1M tokensUS$2.40 in · US$9.60 out | Not clearly above the median | Try |
| 11 | Claude Haiku 4.5 Anthropic | Arena1,415±3 · 129,278 votes | US$2.40 / 1M tokensUS$1.20 in · US$6.00 out | Not clearly above the median | Try |
| 12 | GPT-4.1 OpenAI | Arena1,415±4 · 49,941 votes | US$4.20 / 1M tokensUS$2.40 in · US$9.60 out | Not clearly above the median | Try |
| 13 | OpenAI o4-mini OpenAILegacy | Arena1,391±4 · 44,639 votes | US$2.31 / 1M tokensUS$1.32 in · US$5.28 out | Not clearly above the median | Try |
| 14 | o3-mini OpenAILegacy | Arena1,348±4 · 56,655 votes | US$2.31 / 1M tokensUS$1.32 in · US$5.28 out | Not clearly above the median | Try |
| 15 | GPT-4o OpenAI | Arena1,335±4 · 45,499 votes | US$5.25 / 1M tokensUS$3.00 in · US$12.00 out | Not clearly above the median | Try |
| 16 | GPT-4o Mini OpenAI | Arena1,318±4 · 68,697 votes | US$0.315 / 1M tokensUS$0.18 in · US$0.72 out | Not clearly above the median | Try |
| Without an Arena rating (14), by price | |||||
| – | No Arena rating | US$0.12 / 1M tokensUS$0.060 in · US$0.30 out | Try | ||
| – | GPT-6 Luna OpenAI | No Arena rating | US$0.24 / 1M tokensUS$0.12 in · US$0.60 out | Try | |
| – | No Arena rating | US$0.24 / 1M tokensUS$0.12 in · US$0.60 out | Try | ||
| – | GPT-5.4 Nano OpenAI | No Arena rating | US$0.555 / 1M tokensUS$0.24 in · US$1.50 out | Try | |
| – | No Arena rating | US$0.63 / 1M tokensUS$0.36 in · US$1.44 out | Try | ||
| – | DeepSeek V4.1 Flash DeepSeek | No Arena rating | US$0.63 / 1M tokensUS$0.36 in · US$1.44 out | Try | |
| – | GPT-5 Mini OpenAILegacy | No Arena rating | US$0.825 / 1M tokensUS$0.30 in · US$2.40 out | Try | |
| – | GPT-5.4 Mini OpenAI | No Arena rating | US$2.025 / 1M tokensUS$0.90 in · US$5.40 out | Try | |
| – | DeepSeek V4 Pro DeepSeek | No Arena rating | US$2.376 / 1M tokensUS$1.584 in · US$4.752 out | Try | |
| – | Claude Sonnet 5 Anthropic | No Arena rating | US$4.80 / 1M tokensUS$2.40 in · US$12.00 out | Try | |
| – | GPT-6 Sol OpenAI | No Arena rating | US$4.80 / 1M tokensUS$2.40 in · US$12.00 out | Try | |
| – | Claude Opus 5.5 Anthropic | No Arena rating | US$9.60 / 1M tokensUS$4.80 in · US$24.00 out | Try | |
| – | Claude Fable 5.1 Anthropic | No Arena rating | US$24.00 / 1M tokensUS$12.00 in · US$60.00 out | Try | |
| – | GPT-6 Astra OpenAI | No Arena rating | US$24.00 / 1M tokensUS$12.00 in · US$60.00 out | Try | |
Value = Arena points above the median of the rated models in this list per US dollar of 1M tokens (3:1 input:output); only models whose 95 % interval lies fully above the median get a value rank.
How to read this ranking
- Arena score
- A rating from blind pairwise votes on the public Arena: people compare two anonymous models on the same prompt and pick the better result. Higher is better; ± is the 95 % interval, “votes” the number of battles.
- Price
- The blended price of 1M tokens, 3 parts input to 1 part output, from the same pricing rules that bill your runs (1 credit = $0.01). Input and output prices are listed under it.
- Value
- Value = Arena points above the median of the rated models in this list per US dollar of 1M tokens (3:1 input:output); only models whose 95 % interval lies fully above the median get a value rank.
- Who is listed
- Every model of this category you can run on Railwail right now.
- No Arena rating
- The arena has no entry for exactly this model in the setting we run (for example, only a high-reasoning or a 1080p run is rated). We never estimate a score.
- Rank
- The position within this list, not the rank on the arena.
Source
Quality scores: Text Arena leaderboard, published 13 Sept 2026, retrieved 24 Sept 2026 from the LMArena leaderboard dataset. Used under CC BY 4.0; we show the overall score of the models we can match to a Railwail model and rank them within this list.
Questions about this ranking
Which AI language model has the highest Arena score on Railwail?
Gemini 3.1 Pro by Google DeepMind with an Arena score of 1,487 (Text Arena, published 13 Sept 2026). Next: Gemini 3.1 Pro (1,487) and GPT-5.5 (1,476).
Which AI language model is the cheapest on Railwail?
Granite Code 8B by IBM at US$0.12 / 1M tokens. Prices come from the same pricing rules that bill your runs.
Which AI language model gives the most quality per dollar?
Gemini 3.1 Pro: 45 Arena points above the median of this list (1,442) at US$5.40 / 1M tokens, that is 8 points per dollar.
How is value (quality per dollar) calculated?
Value = Arena points above the median of the rated models in this list per US dollar of 1M tokens (3:1 input:output); only models whose 95 % interval lies fully above the median get a value rank.
Why do some models have no Arena rating?
The public Arena rates many models only in one setting, for example with high reasoning effort, with web search or at 1080p. We show a score only when the rated entry is exactly the model and default setting that runs on Railwail; otherwise the model is listed without a score.