Claude Fable 5.1
Anthropic
Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
claudefablereasoning
$12.00/1M in
$60.00/1M out
Language models for chat, analysis, extraction and agents.
Quick picks
21 models
Anthropic
Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
claudefablereasoning
$12.00/1M in
$60.00/1M out
Anthropic
The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.
claudeopusagentic
$6.00/1M in
$30.00/1M out
Anthropic
Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.
claudeopusagentic
$4.80/1M in
$24.00/1M out
Anthropic
Anthropic's Sonnet model with the best combination of speed and intelligence. 1M-token context window, up to 128K output tokens, adaptive thinking.
claudesonnetbalanced
$2.40/1M in
$12.00/1M out
DeepSeek
DeepSeek's April 2026 flagship. 1.6T MoE / 49B active params, 1M context, rivals top closed-source models on STEM and coding at a fraction of the price.
open-weightsmoecoding
$1.584/1M in
$4.752/1M out
Google DeepMind
Google's latest thinking model. Excels at reasoning, coding, math, and science with massive context window.
reasoningcodingmultimodal
$1.50/1M in
$12.00/1M out
OpenAI
OpenAI's newest flagship model. Improved reasoning, instruction following, and coding over GPT-4o.
codingreasoning
$2.40/1M in
$9.60/1M out
OpenAI
OpenAI's current flagship chat model (released April 2026). Strongest general reasoning, coding and tool use in the GPT-5 line, with vision input and a large context window.
gpt-5reasoningvision
$6.00/1M in
$36.00/1M out
OpenAI
OpenAI's reasoning model optimized for STEM tasks, coding, and math. Uses chain-of-thought reasoning.
reasoningcodingmath
$1.32/1M in
$5.28/1M out
OpenAI
OpenAI's most capable multimodal model. Excellent for complex reasoning, coding, and creative tasks.
fastmultimodal
$3.00/1M in
$12.00/1M out
DeepSeek
Efficiency-optimized variant of DeepSeek V4. 284B MoE / 13B active, 1M context, ultra-low pricing for high-throughput workloads.
open-weightsmoecost-efficient
$0.36/1M in
$1.44/1M out
DeepSeek
DeepSeek's current Flash model (API name deepseek-flash, model version DeepSeek-V4.1-Flash): 1M-token context, up to 384K output tokens, JSON output, tool calls and vision input.
open-weightscost-efficientlong-context
$0.36/1M in
$1.44/1M out
OpenAI
OpenAI's most capable model, built for the hardest end-to-end work. Reasoning, text and image input, function calling and tool use, 1,050,000-token context window and up to 128,000 output tokens. Knowledge cutoff April 30, 2026.
gpt-6reasoningvision
$12.00/1M in
$60.00/1M out
OpenAI
OpenAI's most efficient model for focused, high-volume tasks. Reasoning, text and image input, function calling and tool use, 1,050,000-token context window and up to 128,000 output tokens. Knowledge cutoff May 18, 2026.
gpt-6cost-efficientfast
$0.12/1M in
$0.60/1M out
OpenAI
OpenAI model built to power complex coding and agentic workflows. Reasoning, text and image input, function calling and tool use, 1,050,000-token context window and up to 128,000 output tokens. Knowledge cutoff April 20, 2026.
gpt-6codingagents
$2.40/1M in
$12.00/1M out
OpenAI
Small, fast, and affordable model for lightweight tasks. Great balance of speed and capability.
fastaffordable
$0.18/1M in
$0.72/1M out
OpenAI
Smaller, faster, cheaper member of OpenAI's GPT-5 family. Tuned for high-throughput chat, classification and extraction where the full flagship is overkill.
gpt-5cost-efficientfast
$0.30/1M in
$2.40/1M out
OpenAI
OpenAI GPT-5.1 chat model (November 2025). An earlier GPT-5 point release kept available for compatibility. Good general-purpose reasoning and coding.
gpt-5reasoningvision
$1.50/1M in
$12.00/1M out
OpenAI
OpenAI's o3 reasoning model. Spends compute on a private chain of thought before answering, strong at math, science and hard coding problems that benefit from deliberate reasoning.
o-seriesreasoningstem
$2.40/1M in
$9.60/1M out
OpenAI
OpenAI's o4-mini reasoning model. A cost-efficient reasoning model that trades some depth for much lower price and latency, good for high-volume math and code tasks.
o-seriesreasoningcost-efficient
$1.32/1M in
$5.28/1M out
Community
Meta SeamlessM4T v2 Large. Universal multilingual translation across 100+ languages with text-to-text mode for documents and chat.
translationmetaopen-weights
β $0.0012/run
Anthropic
Anthropic's most powerful model. Exceptional at complex analysis, agentic tasks, and extended reasoning.
reasoningagentic
Currently not offered
Google DeepMind
Google's fastest multimodal model. Supports text, images, audio, and video input.
fastmultimodalaffordable
Currently not offered
Hugging Face
The original Bio_ClinicalBERT from Alsentzer et al., a BERT model initialized from BioBERT and further pretrained on all MIMIC-III clinical notes. Served as a fill-mask endpoint it predicts masked tokens in clinical text and produces clinical embeddings. It is the standard encoder backbone behind many downstream clinical NLP fine-tunes.
medicalresearchnlp
Currently not offered
Hugging Face
Token-classification model from d4data that tags 84 biomedical entity types in clinical and medical text, including disease, sign, symptom, medication, dosage, lab value, body part and procedure. Trained on the Maccrobat clinical case corpus on a DistilBERT base, so it runs cheaply for high-volume tagging.
medicalresearchnlp
Currently not offered
Anthropic
Anthropic's most capable model. Excellent for complex analysis, coding, math, and creative writing.
codinganalysis
Currently not offered
DeepSeek
DeepSeek's refreshed V3.1 release. 671B MoE / 37B active. Tops open-weights leaderboards on coding and reasoning.
open-weightsmoecoding
Currently not offered
xAI
xAI's flagship reasoning model with vision and tool use. 256k context, strong at complex reasoning and STEM tasks.
reasoningvisiontools
Deactivated
xAI's Grok 4.20 reasoning snapshot. Runs an extended thinking pass before answering for multi-step analysis, math and STEM, with a 1M token context window and strong agentic tool calling.
grokreasoninglong-context
Currently not offered
Other
Moonshot AI's 1T-parameter MoE model. Industry-leading agentic coding and tool-use benchmarks.
moonshotkimimoe
Currently not offered
Hugging Face
Token-classification model that extracts 41 medical entity types from clinical text, such as disease, medication, dosage, frequency, lab test, sign and symptom. Fine-tuned on a DeBERTa v3 base, which gives more accurate spans than older BERT-based taggers on the same corpus.
medicalresearchnlp
Currently not offered
MiniMax
MiniMax's 456B hybrid lightning-attention model with native 4M-token context. Industry-leading long-context.
long-contextlightning-attentionopen-weights
Currently not offered
Perplexity
Perplexity's premium web-grounded search model with multi-step reasoning over live sources.
web-searchcitationsrag
Currently not offered
Alibaba (Qwen)
Alibaba's Qwen 3 flagship MoE: 235B total / 22B active. Strong reasoning and tool use, open-weights.
moeopen-weights
Deactivated
DeepSeek
DeepSeek's reasoning model with chain-of-thought capabilities. Excellent for complex problem-solving.
reasoningmath
Currently not offered
xAI
xAI's flagship model. Strong at reasoning, coding, and real-time knowledge with web search capabilities.
reasoningreal-time
Deactivated
AI21 Labs
AI21's flagship hybrid Mamba-Transformer model with a 256k context window for long-document tasks.
long-contextmambahybrid
Currently not offered
AI21 Labs
Cost-efficient hybrid Mamba-Transformer model with 256k context. Tuned for high-throughput RAG.
long-contextmambahybrid
Currently not offered
Hugging Face
BioBERT fine-tuned on the NCBI Disease corpus for disease-name recognition. Given biomedical text it tags spans that mention a disease or condition, using BIO labels. Useful for pulling diagnoses and disease mentions out of abstracts, case reports and clinical notes.
medicalresearchnlp
Currently not offered
Anthropic
Anthropic's fast and affordable model. Great for quick tasks, summarization, and simple coding.
fastaffordable
Currently not offered
Hugging Face
Text-classification model from Betty van Aken that decides whether a medical condition mentioned in a clinical note is present, absent or possible. The target entity is marked in the input with [entity] tags, and the model returns the assertion status. Built on Bio_ClinicalBERT and trained on the 2010 i2b2/VA assertion data.
medicalresearchnlp
Currently not offered
Hugging Face
Token-classification model that tags the three core i2b2 clinical entity types in patient notes: problem, test and treatment. Given a sentence from a discharge summary or progress note it marks which spans are medical problems, which are diagnostic tests and which are treatments or medications.
medicalresearchnlp
Currently not offered
Other
Open-weights multilingual research model from Cohere covering 23 languages. 35B parameters.
coheremultilingualopen-weights
Currently not offered
Cohere's fast lightweight chat model (deprecated Sep 2025). Kept as comparison tombstone.
Currently not offered
Cohere's mid-tier RAG/tool model. Cost-efficient sibling of Command R+ with 128k context.
ragtoolsmultilingual
Currently not offered
Cohere's flagship RAG- and tool-optimized chat model. 128k context, refreshed August 2024.
ragtoolsmultilingual
Currently not offered
xAI's Grok 4.20 standard snapshot. Skips the extended thinking pass for lower-latency answers on tasks that do not need deep deliberation. 1M token context window.
groklong-contextlow-latency
Currently not offered
xAI's Grok 4.20 multi-agent snapshot. Coordinates several specialized agents under one call to handle multi-step workflows that mix research, tool use and synthesis. 1M token context window.
grokmulti-agentagentic
Currently not offered
Meta
Meta's open-source 70B parameter model. Strong all-around performance with multilingual support.
open-source
Currently not offered
Meta
Meta M2M-100 12B many-to-many translation model. Direct translation between 100 languages without pivoting through English.
translationopen-weightsmany-to-many
Deactivated
Google DeepMind
Google MADLAD-400 3B multilingual translation model. 419 languages supported, trained on a 5T-token multilingual corpus with strong low-resource performance.
translationopen-weightsmultilingual
Deactivated
Meta mBART-50 many-to-many translation model. 50 supported languages with strong performance on news and conversational text.
translationopen-weightsmany-to-many
Deactivated
Microsoft
Mixture-of-experts Phi-3.5: 42B total / 6.6B active params. 128k context, multilingual.
open-weightsmoemultilingual
Deactivated
Mistral AI
Mistral's flagship model. Strong reasoning, multilingual, and coding capabilities.
multilingualcoding
Currently not offered
Meta
Meta's No Language Left Behind 3.3B translation model. Direct translation between any pair of 200+ languages including many low-resource African and Asian languages.
translationopen-weightsmultilingual
Deactivated
Meta's distilled 600M NLLB. Same 200-language coverage as the 3B model with a fraction of the parameters, ideal for edge or high-throughput deployment.
translationopen-weightsmultilingual
Deactivated
Nous Research
Full-parameter fine-tune of Llama 3.1 405B by Nous Research. Steerable, uncensored, strong tool use.
open-weightstoolsroleplay
Deactivated
Nous Research
Llama-3.1-70B fine-tune from Nous Research with strong tool/agent capabilities and uncensored alignment.
open-weightstoolsroleplay
Deactivated
Perplexity
Perplexity's fastest and cheapest web-grounded chat model. Live-source citations included.
web-searchcitationscost-efficient
Currently not offered
Perplexity
Perplexity's reasoning model with chain-of-thought and integrated web search.
web-searchreasoning
Currently not offered
Alibaba (Qwen)
Alibaba's powerful open-source model. Excellent at coding, math, and multilingual tasks.
open-sourcecodingmultilingual
Currently not offered
Other
Alibaba's flagship pretrained MoE model. Top-tier reasoning and code performance via DashScope API.
qwenalibabamoe
Currently not offered
Snowflake's open MoE model: 480B total / 17B active params with dense+MoE hybrid architecture.
snowflakemoeopen-weights
Currently not offered
Community
TII's 180B causal decoder chat model fine-tuned on Ultrachat, Platypus and Airoboros.
tiiopen-weights
Deactivated
Community
Unbabel TowerInstruct 13B. Llama-2-based multilingual translation and post-editing model. Strong terminology consistency for enterprise localization.
translationunbabelopen-weights
Deactivated
Other
01.AI's larger general-purpose chat model with 32k context window and strong bilingual performance.
01aichinesebilingual
Currently not offered
Their pages stay online, but they canβt be run at the moment.
Large language models are the workhorse of modern AI: chatbots, agents, summarisers, classifiers, translators. The category is the most crowded on Railwail β OpenAI, Anthropic, Google, Mistral, Meta, DeepSeek, xAI, and dozens of open-weights labs all compete here.
Pricing in text generation is almost always per-token, split into a cheaper input rate and a more expensive output rate. A million input tokens is roughly 750,000 English words, so even verbose prompts rarely push the per-call price above a few cents. Output is where bills grow: agentic loops that re-prompt themselves, long reports, and uncached few-shot examples add up quickly. Cache the system prompt, trim history aggressively, and prefer JSON-mode for structured tasks to keep token counts down.
The trade-off triangle is quality, latency, and cost. Flagship models (GPT-5, Claude 4.6, Gemini 2.5) cost ten to fifty times more than budget tiers (Haiku, Flash, Mini) and respond two to four times slower, but they reason deeper, follow instructions more reliably, and hallucinate less. For high-volume classification or extraction, a budget model is almost always the right call. For long-form analysis, code review, or anything user-facing, the flagship usually pays for itself.
Watch out for context dilution: when you stuff 200K tokens into the window, the model's attention spreads thin and it starts to ignore the middle of the prompt β even on long-context flagships. Retrieve the relevant 8-16K tokens with embeddings instead of pasting the whole document.
Budget models such as GPT-4o mini, GPT-5.4 Nano, DeepSeek V4 Flash and Gemini 3 Flash sit at the low end β from about $0.18 to $0.60 per million input tokens on Railwail today. The ranking changes with every provider price update, so sort the model grid above by price to see the current cheapest option.
In the catalog, Gemini 3.1 Pro lists the largest context window (2M tokens), followed by GPT-5.4, Claude Opus 4.8 and DeepSeek V4 Pro with about 1M tokens each. For most workloads, 128K is more than enough; reach for the long-context tier only when you genuinely need to read a whole codebase or research paper in one prompt.
Open-weights models (Llama 3, Qwen, DeepSeek, Mistral, Mixtral) are catching up fast and win on price-per-token and data sovereignty. Proprietary flagships still lead on reasoning, multilingual coverage, and tool-use reliability. If you're cost-sensitive or need self-hosting, start open. If you're shipping to end-users, start proprietary and optimize down.
GPT-5 leads on raw benchmark scores, math, and code generation; Claude 4.6 leads on long-form writing, instruction-following nuance, and refusal-rate calibration. Both are within 5% of each other on most tasks. Run them side-by-side on your real prompts at /compare/gpt-5-vs-claude-4-6 β the differences are workload-specific.
Railwail is OpenAI-compatible: you only change the `model` parameter. Same endpoint, same SDK, same request body. Try a new model in production by routing 10% of your traffic to it for a week and comparing quality, latency and cost.
Not through the Railwail API yet: the chat endpoint currently does not forward `response_format` or tool definitions to the provider. Ask for JSON in the prompt and validate the answer with a Zod or Pydantic schema; many models follow a clear JSON instruction reliably.
Not yet. The Railwail chat endpoint returns the complete answer in one response; the `stream: true` parameter is currently ignored. For long answers, set a sensible `max_tokens` and show a loading state in your UI.
Railwail runs its platform on servers in Germany and does not train on your prompts. The request itself is processed by the provider of the chosen model, which may be located outside the EU; every model page names that provider, so you can pick one that matches your compliance requirements. Details are on the trust page.
Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.