Best AI language models, ranked by quality

Language models for chat, writing, coding and reasoning, compared on the quality people vote for in the Arena, the price of 1M tokens and quality per dollar.

Best AI language models, ranked by quality

Sorted by Arena score, highest first. Models without an Arena rating follow, by price.

Median Arena score here: 1,442

Best AI language models, ranked by quality
01Arena1,494±4 · 61,128 votesUS$12.00 / 1M tokensUS$6.00 in · US$30.00 out4 pts/$+52 over medianTry
02
Gemini 3.1 Pro
Google DeepMind
Arena1,487±3 · 106,951 votesUS$5.40 / 1M tokensUS$2.40 in · US$14.40 out8 pts/$+45 over medianTry
03
GPT-5.5
OpenAI
Arena1,476±4 · 66,317 votesUS$13.50 / 1M tokensUS$6.00 in · US$36.00 out3 pts/$+34 over medianTry
04
Gemini 3 Flash
Google DeepMindLegacy
Arena1,474±4 · 30,225 votesUS$1.35 / 1M tokensUS$0.60 in · US$3.60 out23 pts/$+31 over medianTry
05Arena1,473±4 · 53,446 votesUS$12.00 / 1M tokensUS$6.00 in · US$30.00 out3 pts/$+31 over medianTry
06Arena1,473±4 · 66,208 votesUS$7.20 / 1M tokensUS$3.60 in · US$18.00 out4 pts/$+30 over medianTry
07
GPT-5.4
OpenAI
Arena1,466±4 · 63,526 votesUS$6.75 / 1M tokensUS$3.00 in · US$18.00 out3 pts/$+23 over medianTry
08
Gemini 2.5 Pro
Google DeepMindLegacy
Arena1,446±3 · 122,554 votesUS$4.125 / 1M tokensUS$1.50 in · US$12.00 out1 pts/$+3 over medianTry
09
GPT-5.1
OpenAI
Arena1,439±4 · 42,982 votesUS$4.125 / 1M tokensUS$1.50 in · US$12.00 outNot clearly above the medianTry
10
OpenAI o3
OpenAILegacy
Arena1,432±4 · 58,579 votesUS$4.20 / 1M tokensUS$2.40 in · US$9.60 outNot clearly above the medianTry
11Arena1,415±3 · 129,278 votesUS$2.40 / 1M tokensUS$1.20 in · US$6.00 outNot clearly above the medianTry
12
GPT-4.1
OpenAI
Arena1,415±4 · 49,941 votesUS$4.20 / 1M tokensUS$2.40 in · US$9.60 outNot clearly above the medianTry
13
OpenAI o4-mini
OpenAILegacy
Arena1,391±4 · 44,639 votesUS$2.31 / 1M tokensUS$1.32 in · US$5.28 outNot clearly above the medianTry
14
o3-mini
OpenAILegacy
Arena1,348±4 · 56,655 votesUS$2.31 / 1M tokensUS$1.32 in · US$5.28 outNot clearly above the medianTry
15
GPT-4o
OpenAI
Arena1,335±4 · 45,499 votesUS$5.25 / 1M tokensUS$3.00 in · US$12.00 outNot clearly above the medianTry
16Arena1,318±4 · 68,697 votesUS$0.315 / 1M tokensUS$0.18 in · US$0.72 outNot clearly above the medianTry
Without an Arena rating (14), by price
–No Arena ratingUS$0.12 / 1M tokensUS$0.060 in · US$0.30 outTry
–No Arena ratingUS$0.24 / 1M tokensUS$0.12 in · US$0.60 outTry
–No Arena ratingUS$0.24 / 1M tokensUS$0.12 in · US$0.60 outTry
–No Arena ratingUS$0.555 / 1M tokensUS$0.24 in · US$1.50 outTry
–No Arena ratingUS$0.63 / 1M tokensUS$0.36 in · US$1.44 outTry
–No Arena ratingUS$0.63 / 1M tokensUS$0.36 in · US$1.44 outTry
–
GPT-5 Mini
OpenAILegacy
No Arena ratingUS$0.825 / 1M tokensUS$0.30 in · US$2.40 outTry
–No Arena ratingUS$2.025 / 1M tokensUS$0.90 in · US$5.40 outTry
–No Arena ratingUS$2.376 / 1M tokensUS$1.584 in · US$4.752 outTry
–No Arena ratingUS$4.80 / 1M tokensUS$2.40 in · US$12.00 outTry
–
GPT-6 Sol
OpenAI
No Arena ratingUS$4.80 / 1M tokensUS$2.40 in · US$12.00 outTry
–No Arena ratingUS$9.60 / 1M tokensUS$4.80 in · US$24.00 outTry
–No Arena ratingUS$24.00 / 1M tokensUS$12.00 in · US$60.00 outTry
–No Arena ratingUS$24.00 / 1M tokensUS$12.00 in · US$60.00 outTry

Value = Arena points above the median of the rated models in this list per US dollar of 1M tokens (3:1 input:output); only models whose 95 % interval lies fully above the median get a value rank.

How to read this ranking

Arena score
A rating from blind pairwise votes on the public Arena: people compare two anonymous models on the same prompt and pick the better result. Higher is better; ± is the 95 % interval, “votes” the number of battles.
Price
The blended price of 1M tokens, 3 parts input to 1 part output, from the same pricing rules that bill your runs (1 credit = $0.01). Input and output prices are listed under it.
Value
Value = Arena points above the median of the rated models in this list per US dollar of 1M tokens (3:1 input:output); only models whose 95 % interval lies fully above the median get a value rank.
Who is listed
Every model of this category you can run on Railwail right now.
No Arena rating
The arena has no entry for exactly this model in the setting we run (for example, only a high-reasoning or a 1080p run is rated). We never estimate a score.
Rank
The position within this list, not the rank on the arena.

Source

Quality scores: Text Arena leaderboard, published Sep 13, 2026, retrieved Sep 24, 2026 from the LMArena leaderboard dataset. Used under CC BY 4.0; we show the overall score of the models we can match to a Railwail model and rank them within this list.

arena.ai leaderboardDatasetCC BY 4.0

Questions about this ranking

Which AI language model has the highest Arena score on Railwail?

Gemini 3.1 Pro by Google DeepMind with an Arena score of 1,487 (Text Arena, published Sep 13, 2026). Next: Gemini 3.1 Pro (1,487) and GPT-5.5 (1,476).

Which AI language model is the cheapest on Railwail?

Granite Code 8B by IBM at US$0.12 / 1M tokens. Prices come from the same pricing rules that bill your runs.

Which AI language model gives the most quality per dollar?

Gemini 3.1 Pro: 45 Arena points above the median of this list (1,442) at US$5.40 / 1M tokens, that is 8 points per dollar.

How is value (quality per dollar) calculated?

Value = Arena points above the median of the rated models in this list per US dollar of 1M tokens (3:1 input:output); only models whose 95 % interval lies fully above the median get a value rank.

Why do some models have no Arena rating?

The public Arena rates many models only in one setting, for example with high reasoning effort, with web search or at 1080p. We show a score only when the rated entry is exactly the model and default setting that runs on Railwail; otherwise the model is listed without a score.

Best AI language models ranked by quality, price and value | Railwail