Text & Chat Models

Language models for chat, analysis, extraction and agents.

Available
21of 66
Providers
4
Price range
$0.12 – $12.00per 1M input tokens

21 models

  • New

    Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    claudefablereasoning

    Modalities: Text, Image1M context

    $12.00/1M in

    $60.00/1M out

  • New

    The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    claudeopusagentic

    Modalities: Text, Image1M context

    $6.00/1M in

    $30.00/1M out

  • New

    Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    claudeopusagentic

    Modalities: Text, Image1M context

    $4.80/1M in

    $24.00/1M out

  • New

    Anthropic's Sonnet model with the best combination of speed and intelligence. 1M-token context window, up to 128K output tokens, adaptive thinking.

    claudesonnetbalanced

    Modalities: Text, Image1M context

    $2.40/1M in

    $12.00/1M out

  • New

    DeepSeek's April 2026 flagship. 1.6T MoE / 49B active params, 1M context, rivals top closed-source models on STEM and coding at a fraction of the price.

    open-weightsmoecoding

    Modalities: Text1M context

    $1.584/1M in

    $4.752/1M out

  • Gemini 2.5 Pro

    Google DeepMind

    New

    Google's latest thinking model. Excels at reasoning, coding, math, and science with massive context window.

    reasoningcodingmultimodal

    Modalities: Text1M context

    $1.50/1M in

    $12.00/1M out

  • GPT-4.1

    OpenAI

    New

    OpenAI's newest flagship model. Improved reasoning, instruction following, and coding over GPT-4o.

    codingreasoning

    Modalities: Text1M context

    $2.40/1M in

    $9.60/1M out

  • GPT-5.5

    OpenAI

    New

    OpenAI's current flagship chat model (released April 2026). Strongest general reasoning, coding and tool use in the GPT-5 line, with vision input and a large context window.

    gpt-5reasoningvision

    Modalities: Text, Image400K context

    $6.00/1M in

    $36.00/1M out

  • o3-mini

    OpenAI

    New

    OpenAI's reasoning model optimized for STEM tasks, coding, and math. Uses chain-of-thought reasoning.

    reasoningcodingmath

    Modalities: Text200K context

    $1.32/1M in

    $5.28/1M out

  • GPT-4o

    OpenAI

    OpenAI's most capable multimodal model. Excellent for complex reasoning, coding, and creative tasks.

    fastmultimodal

    Modalities: Text128K context

    $3.00/1M in

    $12.00/1M out

  • New

    Efficiency-optimized variant of DeepSeek V4. 284B MoE / 13B active, 1M context, ultra-low pricing for high-throughput workloads.

    open-weightsmoecost-efficient

    Modalities: Text1M context

    $0.36/1M in

    $1.44/1M out

  • New

    DeepSeek's current Flash model (API name deepseek-flash, model version DeepSeek-V4.1-Flash): 1M-token context, up to 384K output tokens, JSON output, tool calls and vision input.

    open-weightscost-efficientlong-context

    Modalities: Text, Image1M context

    $0.36/1M in

    $1.44/1M out

  • New

    OpenAI's most capable model, built for the hardest end-to-end work. Reasoning, text and image input, function calling and tool use, 1,050,000-token context window and up to 128,000 output tokens. Knowledge cutoff April 30, 2026.

    gpt-6reasoningvision

    Modalities: Text1.1M context

    $12.00/1M in

    $60.00/1M out

  • New

    OpenAI's most efficient model for focused, high-volume tasks. Reasoning, text and image input, function calling and tool use, 1,050,000-token context window and up to 128,000 output tokens. Knowledge cutoff May 18, 2026.

    gpt-6cost-efficientfast

    Modalities: Text1.1M context

    $0.12/1M in

    $0.60/1M out

  • GPT-6 Sol

    OpenAI

    New

    OpenAI model built to power complex coding and agentic workflows. Reasoning, text and image input, function calling and tool use, 1,050,000-token context window and up to 128,000 output tokens. Knowledge cutoff April 20, 2026.

    gpt-6codingagents

    Modalities: Text1.1M context

    $2.40/1M in

    $12.00/1M out

  • Small, fast, and affordable model for lightweight tasks. Great balance of speed and capability.

    fastaffordable

    Modalities: Text128K context

    $0.18/1M in

    $0.72/1M out

  • Smaller, faster, cheaper member of OpenAI's GPT-5 family. Tuned for high-throughput chat, classification and extraction where the full flagship is overkill.

    gpt-5cost-efficientfast

    Modalities: Text, Image400K context

    $0.30/1M in

    $2.40/1M out

  • GPT-5.1

    OpenAI

    OpenAI GPT-5.1 chat model (November 2025). An earlier GPT-5 point release kept available for compatibility. Good general-purpose reasoning and coding.

    gpt-5reasoningvision

    Modalities: Text, Image400K context

    $1.50/1M in

    $12.00/1M out

  • OpenAI o3

    OpenAI

    OpenAI's o3 reasoning model. Spends compute on a private chain of thought before answering, strong at math, science and hard coding problems that benefit from deliberate reasoning.

    o-seriesreasoningstem

    Modalities: Text, Image200K context

    $2.40/1M in

    $9.60/1M out

  • OpenAI's o4-mini reasoning model. A cost-efficient reasoning model that trades some depth for much lower price and latency, good for high-volume math and code tasks.

    o-seriesreasoningcost-efficient

    Modalities: Text, Image200K context

    $1.32/1M in

    $5.28/1M out

  • Meta SeamlessM4T v2 Large. Universal multilingual translation across 100+ languages with text-to-text mode for documents and chat.

    translationmetaopen-weights

    Modalities: Text4.1K context

    β‰ˆ $0.0012/run

42 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Text and chat models for production AI workloads

Large language models are the workhorse of modern AI: chatbots, agents, summarisers, classifiers, translators. The category is the most crowded on Railwail β€” OpenAI, Anthropic, Google, Mistral, Meta, DeepSeek, xAI, and dozens of open-weights labs all compete here.

Pricing, trade-offs and pitfalls

Pricing in text generation is almost always per-token, split into a cheaper input rate and a more expensive output rate. A million input tokens is roughly 750,000 English words, so even verbose prompts rarely push the per-call price above a few cents. Output is where bills grow: agentic loops that re-prompt themselves, long reports, and uncached few-shot examples add up quickly. Cache the system prompt, trim history aggressively, and prefer JSON-mode for structured tasks to keep token counts down.

The trade-off triangle is quality, latency, and cost. Flagship models (GPT-5, Claude 4.6, Gemini 2.5) cost ten to fifty times more than budget tiers (Haiku, Flash, Mini) and respond two to four times slower, but they reason deeper, follow instructions more reliably, and hallucinate less. For high-volume classification or extraction, a budget model is almost always the right call. For long-form analysis, code review, or anything user-facing, the flagship usually pays for itself.

Watch out for context dilution: when you stuff 200K tokens into the window, the model's attention spreads thin and it starts to ignore the middle of the prompt β€” even on long-context flagships. Retrieve the relevant 8-16K tokens with embeddings instead of pasting the whole document.

Typical tasks

  • Customer support chatbots
  • Long-document summarization
  • Code generation in IDEs
  • Retrieval-augmented question answering
  • Multilingual translation pipelines
  • Agentic workflows and tool-use

Guides by use case

Model comparisons

Frequently asked questions

Which LLM is the cheapest on Railwail?

Budget models such as GPT-4o mini, GPT-5.4 Nano, DeepSeek V4 Flash and Gemini 3 Flash sit at the low end β€” from about $0.18 to $0.60 per million input tokens on Railwail today. The ranking changes with every provider price update, so sort the model grid above by price to see the current cheapest option.

Which model has the longest context window?

In the catalog, Gemini 3.1 Pro lists the largest context window (2M tokens), followed by GPT-5.4, Claude Opus 4.8 and DeepSeek V4 Pro with about 1M tokens each. For most workloads, 128K is more than enough; reach for the long-context tier only when you genuinely need to read a whole codebase or research paper in one prompt.

Open-source vs proprietary β€” which should I pick?

Open-weights models (Llama 3, Qwen, DeepSeek, Mistral, Mixtral) are catching up fast and win on price-per-token and data sovereignty. Proprietary flagships still lead on reasoning, multilingual coverage, and tool-use reliability. If you're cost-sensitive or need self-hosting, start open. If you're shipping to end-users, start proprietary and optimize down.

GPT-5 vs Claude 4.6 β€” which is better?

GPT-5 leads on raw benchmark scores, math, and code generation; Claude 4.6 leads on long-form writing, instruction-following nuance, and refusal-rate calibration. Both are within 5% of each other on most tasks. Run them side-by-side on your real prompts at /compare/gpt-5-vs-claude-4-6 β€” the differences are workload-specific.

How do I switch models in my code?

Railwail is OpenAI-compatible: you only change the `model` parameter. Same endpoint, same SDK, same request body. Try a new model in production by routing 10% of your traffic to it for a week and comparing quality, latency and cost.

Is JSON mode supported?

Not through the Railwail API yet: the chat endpoint currently does not forward `response_format` or tool definitions to the provider. Ask for JSON in the prompt and validate the answer with a Zod or Pydantic schema; many models follow a clear JSON instruction reliably.

Does Railwail support streaming?

Not yet. The Railwail chat endpoint returns the complete answer in one response; the `stream: true` parameter is currently ignored. For long answers, set a sensible `max_tokens` and show a loading state in your UI.

Is the API GDPR-compliant?

Railwail runs its platform on servers in Germany and does not train on your prompts. The request itself is processed by the provider of the chosen model, which may be located outside the EU; every model page names that provider, so you can pick one that matches your compliance requirements. Details are on the trust page.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.