AI Model Leaderboard

Every model you can run on Railwail right now, by category: quality from the public Arena leaderboard, the live price of a default call, the context window and run times measured on our own jobs. Scores only where the arena rates exactly the model we run.

Arena scores published Sep 22, 2026Run times from Railwail jobs, last 90 days

Models you can run
274
Labs
34
Categories
9

Today’s record holders

  1. Highest Arena score · Text

    Arena 1,487

  2. Highest Arena score · Image

    Arena 1,255

  3. Highest Arena score · Video

    Arena 1,362

  4. Lowest LLM input price

    $0.060/1M in

    $0.30/1M out

  5. Largest context window

    1.1M tokens

  6. Lowest price per image

    $0.0036/image

  7. Lowest price per video second

    $0.0432/s

  8. Lowest text-to-speech price

    $0.018/1k chars

  9. Lowest transcription run

    ≈ $0.0034/run

Best models by category

Quality from the public Arena, prices from our pricing rules. Open a category to sort by quality, price or value.

Filter by labTap a lab to filter every board.

01

Language models

30

Chat, reasoning and code models billed per token. Sorted by Arena score, highest first; models without a rating follow by price.

Claude Opus 4.7
Arena 1494 · $30.00/1M out · 1M ctx
1494$6.00/1M in$6.00$30.001M·Try
Gemini 3.1 Pro
Google DeepMind
Arena 1487 · $14.40/1M out · 1M ctx
1487$2.40/1M in$2.40$14.401M·Try
GPT-5.5
OpenAI
Arena 1476 · $36.00/1M out · 400K ctx
1476$6.00/1M in$6.00$36.00400K·Try
Gemini 3 Flash
Google DeepMindLegacy
Arena 1474 · $3.60/1M out · 1M ctx
1474$0.60/1M in$0.60$3.601M·Try
Claude Opus 4.8
Arena 1473 · $30.00/1M out · 1M ctx · median 3.2 s
1473$6.00/1M in$6.00$30.001M3.2 s198 runsTry
Claude Sonnet 4.6
Arena 1473 · $18.00/1M out · 1M ctx
1473$3.60/1M in$3.60$18.001M·Try
GPT-5.4
OpenAI
Arena 1466 · $18.00/1M out · 1.1M ctx
1466$3.00/1M in$3.00$18.001.1M·Try
Gemini 2.5 Pro
Google DeepMindLegacy
Arena 1446 · $12.00/1M out · 1M ctx
1446$1.50/1M in$1.50$12.001M·Try
GPT-5.1
OpenAI
Arena 1439 · $12.00/1M out · 400K ctx
1439$1.50/1M in$1.50$12.00400K·Try
OpenAI o3
OpenAILegacy
Arena 1432 · $9.60/1M out · 200K ctx
1432$2.40/1M in$2.40$9.60200K·Try
Claude Haiku 4.5 Anthropic$1.20/1M in
GPT-4.1 OpenAI$2.40/1M in
OpenAI o4-mini OpenAI$1.32/1M in
o3-mini OpenAI$1.32/1M in
GPT-4o OpenAI$3.00/1M in
GPT-4o Mini OpenAI$0.18/1M in
Granite Code 8B IBM$0.060/1M in
GPT-6 Luna OpenAI$0.12/1M in
Granite Code 20B IBM$0.12/1M in
GPT-5.4 Nano OpenAI$0.24/1M in
GPT-5 Mini OpenAI$0.30/1M in
DeepSeek V4 Flash DeepSeek$0.36/1M in
DeepSeek V4.1 Flash DeepSeek$0.36/1M in
GPT-5.4 Mini OpenAI$0.90/1M in
DeepSeek V4 Pro DeepSeek$1.584/1M in
Claude Sonnet 5 Anthropic$2.40/1M in
GPT-6 Sol OpenAI$2.40/1M in
Claude Opus 5.5 Anthropic$4.80/1M in
Claude Fable 5.1 Anthropic$12.00/1M in
GPT-6 Astra OpenAI$12.00/1M in
Median run “·” = Fewer than 5 completed runs in 90 days

Arena scores: Text Arena leaderboard by LMArena, published Sep 13, 2026, CC BY 4.0.Ranked by quality, price and value →

A usage column appears once 5 models have runs from 3 or more different accounts in 30 days.

02

Image generation & editing

83/134

Text-to-image, editing, upscaling and background removal. Sorted by Arena score, highest first; models without a rating follow by price.

Qwen Image 3 Pro
Alibaba (Qwen)
Arena 1255
1255$0.048/image·Try
Nano Banana 2 Lite
Google DeepMind
Arena 1250
1250$0.0408/image·Try
Ideogram V4 Quality
Ideogram
Arena 1204
1204$0.12/image·Try
Grok Imagine Image
xAI
Arena 1171
1171$0.024/image·Try
Recraft V4.1 Utility Pro
Recraft
Arena 1169
1169$0.30/image·Try
Hunyuan Image 3
Tencent
Arena 1151
1151$0.096/image·Try
Nano Banana
Google DeepMind
Arena 1150
1150$0.0468/image·Try
Recraft V4.1 Pro
Recraft
Arena 1131
1131$0.30/image·Try
Google Imagen 4
Google DeepMind
Arena 1129
1129$0.048/image·Try
Qwen Image 2512
Alibaba (Qwen)
Arena 1125
1125$0.024/image·Try
Recraft V4 Recraft$0.048/image
Flux Kontext Max Black Forest Labs$0.096/image
Flux Kontext Pro Black Forest Labs$0.048/image
Qwen Image Alibaba (Qwen)$0.030/image
Ideogram v3 Quality Ideogram$0.108/image
Photon Luma AI$0.036/image
P Image Pruna AI$0.0060/image
Gen4 Image Runway$0.096/image
Recraft V3 Recraft$0.048/image
FLUX 1.1 Pro Black Forest Labs$0.048/image
Ideogram v2 Ideogram$0.096/image
Stable Diffusion 3.5 Large Stability AI$0.078/image
851-Labs Background Remover Community≈ $0.00050/run
Remove Background (lucataco) Community≈ $0.00060/run
Remove Object (LaMa) Community≈ $0.00090/run
DDColor Community≈ $0.0012/run
Html To Image Community$0.0012/image
Stable Diffusion XL Stability AI≈ $0.0015/run
Vectorizer (VTracer) Community≈ $0.0022/run
Real-ESRGAN 4x Community$0.0024/image
SDXL Inpainting Community≈ $0.0026/run
GFPGAN v1.4 Tencent ARC≈ $0.0027/run
BiRefNet Background Removal Community≈ $0.0036/run
FLUX.1 [schnell] Black Forest Labs$0.0036/image
FLUX.1 Redux Black Forest Labs$0.0036/image
Consistent Character Community≈ $0.0038/run
InstructPix2Pix Community≈ $0.0039/run
CodeFormer Community≈ $0.0041/run
Rembg Community≈ $0.0047/run
Real-ESRGAN Anime 4x Community≈ $0.0051/run
Sticker Maker Community≈ $0.0055/run
Flux Fast Community$0.0060/image
Hidream L1 Fast Community$0.0060/image
PhotoMaker Tencent ARC≈ $0.0080/run
Swin2SR Community≈ $0.0088/run
Icons (SDXL Flat Pop) Community≈ $0.0094/run
Bringing Old Photos Back to Life Microsoft≈ $0.010/run
Face to Many Community≈ $0.0105/run
Image 01 MiniMax$0.012/image
Photon Flash Luma AI$0.012/image
Recraft Vectorize Recraft$0.012/image
ControlNet Depth Community≈ $0.0144/run
SDXL Emoji Community≈ $0.0156/run
Janus Pro 7B DeepSeek≈ $0.0169/run
FLUX PuLID Community≈ $0.018/run
Lucid Origin Community$0.0201/image
BRIA Remove Background Bria AI$0.0216/image
Clarity Upscaler Community≈ $0.024/run
Hunyuan Image 2.1 Tencent$0.024/image
Wan 2.2 Image Community$0.024/image
Recraft 20B Recraft$0.0264/image
IDM-VTON (Virtual Try-On) Community≈ $0.0277/run
IP-Adapter FaceID Plus v2 Community≈ $0.0277/run
Face to Sticker Community≈ $0.0289/run
Flux Krea Dev Black Forest Labs$0.030/image
FLUX.1 [dev] Black Forest Labs$0.030/image
FLUX.1 Canny [dev] Black Forest Labs$0.030/image
FLUX.1 Depth [dev] Black Forest Labs$0.030/image
FLUX.1 Kontext [dev] Black Forest Labs$0.030/image
Ideogram V2A Turbo Ideogram$0.030/image
Ideogram v3 Turbo Ideogram$0.036/image
Qwen Image 3 Alibaba (Qwen)$0.036/image
Qwen-Image-Edit Alibaba (Qwen)$0.036/image
Seedream 4 ByteDance$0.036/image
Wan 2.7 Image Alibaba (Wan)$0.036/image
Wan 2.7 Image Pro Alibaba (Wan)$0.036/image
LogoAI (SDXL Logo Generator) Community≈ $0.0372/run
Flux Dev Black Forest Labs$0.0384/image
Qwen Image 2 Alibaba (Qwen)$0.042/image
Seedream 5 Lite ByteDance$0.042/image
Stable Diffusion 3 Stability AI$0.042/image
Stable Diffusion 3.5 Medium Stability AI$0.042/image
TRELLIS (3D) Community≈ $0.042/run
ESRGAN Classic Community≈ $0.0468/run
Fibo Bria AI$0.048/image
FLUX.1 Fill [dev] Black Forest Labs$0.048/image
Grok Imagine Image 2 xAI$0.048/image
Ideogram V2A Ideogram$0.048/image
Image 3.2 Bria AI$0.048/image
Professional Headshot (FLUX Kontext) Black Forest Labs$0.048/image
Recraft V4.1 Recraft$0.048/image
Recraft V4.1 Utility Recraft$0.048/image
Seedream 4.5 ByteDance$0.048/image
Stable Diffusion 3.5 Large Turbo Stability AI$0.048/image
Recraft 20B SVG Recraft$0.0528/image
Magnific-Style Upscaler Community≈ $0.0589/run
FLUX.1 Canny Black Forest Labs$0.060/image
FLUX.1 Depth Black Forest Labs$0.060/image
FLUX.1 Fill Black Forest Labs$0.060/image
Ideogram V2 Turbo Ideogram$0.060/image
AuraFlow v0.3 Community≈ $0.0648/run
Flux Pro Black Forest Labs$0.066/image
Ad Inpaint (Product Photo) Community≈ $0.0672/run
ControlNet Canny Community≈ $0.0697/run
Point-E Community≈ $0.0708/run
Flux 1.1 Pro Ultra Black Forest Labs$0.072/image
Ideogram V3 Balanced Ideogram$0.072/image
Ideogram V4 Balanced Ideogram$0.072/image
Playground v2.5 (1024px Aesthetic) Playground AI≈ $0.0732/run
IC-Light (Product Relighting) Community≈ $0.0781/run
Nano Banana 2 Google DeepMind$0.0804/image
InstantID Community≈ $0.084/run
Rd Animation Community$0.084/image
Kuaishou Kolors Community≈ $0.0889/run
FLUX.1-dev Inpainting Community≈ $0.090/run
Qwen Image 2 Pro Alibaba (Qwen)$0.090/image
Cartoonify Community≈ $0.096/run
Recraft v3 SVG Recraft$0.096/image
Recraft V4 SVG Recraft$0.096/image
Shap-E (OpenAI) Community≈ $0.1116/run
Gpt Image 2 OpenAI$0.1536/image
Hunyuan3D 2.0 Tencent≈ $0.156/run
Gpt Image 1.5 OpenAI$0.1632/image
Nano Banana Pro Google DeepMind$0.18/image
OOTDiffusion (Try-On) Community≈ $0.1801/run
Hunyuan3D 2.1 Community≈ $0.264/run
CCSR (Content-Consistent SR) Community≈ $0.2761/run
Gpt Image 2.5 Flare OpenAI$0.30/image
Gpt Image 2.5 Sunburst OpenAI$0.30/image
Recraft V4 Pro Recraft$0.30/image
SUPIR Upscaler Community≈ $0.30/run
SUPIR Community≈ $0.4921/run
StarVector 8B (image-to-SVG) Community$0.00117/GPU s
DreamGaussian Community$0.00168/GPU s
Median run “·” = Fewer than 5 completed runs in 90 days

Arena scores: Text-to-Image Arena leaderboard by LMArena, published Sep 22, 2026, CC BY 4.0.Ranked by quality, price and value →

03

Video generation

31/46

Text-to-video and image-to-video. Sorted by Arena score, highest first; models without a rating follow by price.

Google Veo 3.1
Google DeepMind
Arena 1362
1362$0.48/s$1.92 for 4 s·Try
Google Veo 3.1 Fast
Google DeepMind
Arena 1358
1358$0.18/s$0.72 for 4 s·Try
Grok Imagine Video
xAI
Arena 1342 · median 48 s
1342$0.060/s$0.30 for 5 s48 s31 runsTry
PixVerse v5.6
PixVerse
Arena 1240
1240$0.084/s$0.42 for 5 s·Try
Runway Gen 4.5
Runway
Arena 1225
1225$0.144/s$0.72 for 5 s·Try
Hailuo 2.3
MiniMax
Arena 1206
1206$0.336/video·Try
Seedance Pro
ByteDance
Arena 1191
1191$0.18/s$0.90 for 5 s·Try
1181$0.324/video·Try
Google Veo 2
Google DeepMindLegacySuccessor: Google Veo 3.1
Arena 1164
1164$0.60/s$4.80 for 8 s·Try
Kling v2.1 Master
Kuaishou (Kling)LegacySuccessor: Kling v3
Arena 1163
1163$0.336/s$1.68 for 5 s·Try
Seedance Lite ByteDance$0.0432/s
Luma Ray-2 720p Luma AI$0.216/s
Mochi 1 Genmo≈ $0.5041/run
FILM Frame Interpolation Google Research≈ $0.0016/run
Wav2Lip Community≈ $0.0067/run
LTX-Video (Lightricks) Lightricks≈ $0.0229/run
SwinIR Video Community≈ $0.0277/run
Wan 2.2 5B Fast Alibaba (Wan)$0.030/video
RIFE Frame Interpolation Community≈ $0.0444/run
Wan 2.2 Image-to-Video Alibaba (Wan)$0.060/video
Wan 2.2 Text-to-Video Alibaba (Wan)$0.060/video
MuseTalk Community≈ $0.0624/run
ToonCrafter Community≈ $0.0852/run
LivePortrait Community≈ $0.0949/run
SadTalker Community≈ $0.1165/run
AnimateDiff Community≈ $0.1201/run
VideoCrafter Community≈ $0.1561/run
Kling v2.1 Kuaishou (Kling)$0.060/s
Runway Gen-4 Turbo Runway$0.060/s
Luma Ray Flash 2 Luma AI$0.072/s
Seedance 1 Pro Fast ByteDance$0.072/s
CogVideoX-5B (open) Zhipu AI≈ $0.4081/run
MagicAnimate Community≈ $0.4081/run
Kling V2.5 Turbo Pro Kuaishou (Kling)$0.084/s
EchoMimic Community≈ $0.48/run
Mochi 1 Community≈ $0.5041/run
CogVideoX-5B Community≈ $0.60/run
Minimax Video MiniMax$0.60/video
Wan 3 Alibaba (Qwen)$0.12/s
Google Veo 3 Fast Google DeepMind$0.18/s
Kling v3 Kuaishou (Kling)$0.2688/s
Kling v3 Omni Kuaishou (Kling)$0.2688/s
HunyuanVideo Tencent≈ $3.06/run
Google Veo 3 (Replicate) Google DeepMind$0.48/s
AnimateDiff Lightning Community$0.00117/GPU s
V-Express Community$0.00168/GPU s
Median run “·” = Fewer than 5 completed runs in 90 days

Arena scores: Text-to-Video Arena leaderboard by LMArena, published Sep 22, 2026, CC BY 4.0.Ranked by quality, price and value →

04

Text to speech

7/14

Voices from text, including voice cloning. Sorted by the price of one default call, lowest first.

$0.018/1k charsTry
Qwen3 TTS
Alibaba (Qwen)
$0.024/1k charsTry
Chatterbox
Resemble AI
$0.030/1k charsTry
$0.036/1k charsTry
Kokoro TTS 82M Community≈ $0.00030/run
StyleTTS 2 Community≈ $0.00040/run
Parler-TTS Community≈ $0.0024/run
Spark TTS Community≈ $0.0035/run
AudioLDM 2
Haohe Liu
≈ $0.0157/runTry
RVC Voice Conversion Community≈ $0.0505/run
Riffusion
Riffusion
≈ $0.0576/runTry
OpenVoice v2 Community≈ $0.0673/run
Tortoise TTS Community≈ $0.0816/run
≈ $0.0972/runTry

Ranked by quality, price and value →

05

Speech to text

5

Transcription and speaker diarization. Sorted by the price of one default call, lowest first.

Whisper
OpenAI
≈ $0.0034/runTry
≈ $0.0056/runTry
≈ $0.0063/runTry
WhisperX
Community
≈ $0.0241/runTry
SeamlessM4T
Community
≈ $0.156/runTry

Ranked by quality, price and value →

06

Vision & multimodal

30

Open models that caption, read or segment images, billed by GPU time. Sorted by the price of one default call, lowest first.

BLIP
Salesforce
≈ $0.00030/runTry
MiDaS v3.1
Community
≈ $0.00030/runTry
Segformer B5
Community
≈ $0.00030/runTry
ZoeDepth
Community
≈ $0.00060/runTry
≈ $0.0012/runTry
GOT-OCR 2.0
Community
≈ $0.0012/runTry
Idefics3 8B
Community
≈ $0.0012/runTry
OpenPose
Community
≈ $0.0012/runTry
EasyOCR
Community
≈ $0.0014/runTry
Grounded-SAM
Community
≈ $0.0015/runTry
MiniCPM-V 2.6 Community≈ $0.0015/run
Moondream2 Community≈ $0.0020/run
Qwen2-VL 7B Instruct Community≈ $0.0024/run
Llama 3.2 Vision 11B (Ollama) Community≈ $0.0039/run
Depth Anything v2 Community≈ $0.0050/run
Llama 3.2 Vision 90B Community≈ $0.0070/run
DeepSeek-VL 7B DeepSeek≈ $0.0086/run
GLPN Depth Community≈ $0.0094/run
CogVLM2 19B Community≈ $0.0114/run
Dots OCR Community≈ $0.0132/run
SAM 2 (Segment Anything 2) Meta≈ $0.018/run
olmOCR Community≈ $0.0205/run
PaddleOCR v3 Community≈ $0.0216/run
Mask2Former Meta≈ $0.0276/run
Lotus-G Community≈ $0.0408/run
CLIP Interrogator Community≈ $0.0457/run
Molmo 7B Community≈ $0.0505/run
Donut Document Community≈ $0.0648/run
Marigold Community≈ $0.0936/run
LLaVA 1.6 Vicuna 13B Community≈ $0.1021/run
07

Open-weight text & code

10

Open models on dedicated GPUs, billed by GPU time. Sorted by the price of one default call, lowest first.

≈ $0.0012/runTry
≈ $0.0030/runTry
≈ $0.0065/runTry
≈ $0.0074/runTry
≈ $0.0253/runTry
≈ $0.0408/runTry
≈ $0.0504/runTry
≈ $0.0924/runTry
≈ $0.2161/runTry
≈ $1.188/runTry
08

Music & audio

3

Music and sound generation. Sorted by the price of one default call, lowest first.

MAGNeT
Community
≈ $0.0030/runTry
≈ $0.0069/runTry
≈ $0.0529/runTry

Ranked by quality, price and value →

09

Embeddings

2

Vectors for search and retrieval, billed per token. Sorted by input price per 1M tokens, lowest first.

$0.024/1M in8.2KTry
$0.156/1M in8.2KTry
10

How this leaderboard works

Every figure comes from our catalog, our pricing rules, our own job records or, for quality, the public Arena leaderboard. What we cannot source, we leave out.

What is listed
Only models you can run right now: active and with a verified price. Left out: 130 catalog entries without a price (they cannot run), 0 removed models and 5 duplicate entries of the same provider model. Legacy models stay listed and link their successor.
Price
What one call with the model’s default settings costs on Railwail, from the same pricing rules that bill your runs (1 credit = $0.01). Token models show the price per 1M input and output tokens. Prices marked ≈ are billed by GPU time: the figure is a typical run, the charge follows the real run time.
Median run time
Measured on completed Railwail jobs: the median of the newest up to 500 runs in the last 90 days, shown from 5 runs. It includes queueing at the provider. For language models it depends on the answer length, so the LLM board is never sorted by speed. Measured today: 5 models.
Quality
Arena scores from the public Arena leaderboard (formerly LMArena, CC BY 4.0), where people compare two anonymous models on the same prompt and vote. A model gets a score only when the arena entry is exactly the model and default setting we run (reviewed list, 56 models today). No score is estimated.
Context window
The maximum tokens per request as listed in our catalog. Not re-measured by us.
Usage
Runs per model are only shown once they say something: at least 5 models with runs from 3 or more different accounts in 30 days. Until then there is no popularity ranking.
What we don’t show
No votes of our own and no estimated scores: our own arena has 0 decided battles and 2 model votes, far too few for a rating. Models the public Arena has not rated in the setting we run show “No Arena rating”.

Source:arena.ai leaderboardDatasetCC BY 4.0

Found a wrong price or a missing model? Write to us

Questions about the leaderboard

Which language model is the cheapest on Railwail?

Right now Granite Code 8B by IBM at $0.060/1M in (output $0.30/1M out). The language model board lists all 30 language models by input price.

Which model has the largest context window?

GPT-5.4 by OpenAI with 1.1M tokens per request, as listed in our catalog.

What is the lowest price per image?

FLUX.1 [schnell] by Black Forest Labs at $0.0036/image. Models billed by GPU time can cost less per run; their price, marked ≈, is a typical run and the charge follows the real run time.

Where do the quality scores come from?

From the public Arena leaderboard (formerly LMArena): people compare two anonymous models on the same prompt and vote for the better result; the votes give each model a score. We use the published leaderboards (text: Sep 13, 2026, image and video: Sep 22, 2026) under CC BY 4.0 and show a score only where the arena rates exactly the model and default setting that runs on Railwail.

How current are the prices?

They come from the same pricing rules that bill your runs and were last checked on Sep 24, 2026. The snapshot on this page is refreshed about once an hour.

How do I run one of these models?

Open the model and try it on its page, or create an account, get an API key and call it through the Railwail API at https://railwail.com/api/v1. Your balance is prepaid in USD; runs are billed by the same pricing rules as the prices shown here.

AI Model Leaderboard: quality, price and value by category | Railwail