AI Models

One API key for every available model: compare prices in USD, filter by provider and modality, then try a model in the playground.

Available
274of 490
Providers
34
Price range
$0.00030 – $4.80per default run

273 models

  • Claude Fable 5.1

    AnthropicText & Chat

    New

    Anthropic's model for the most demanding reasoning and long-horizon agentic work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    claudefablereasoning

    Modalities: Text, Image1M context

    $12.00/1M in

    $60.00/1M out

  • Claude Opus 4.7

    AnthropicMultimodal

    New

    Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    reasoningagenticvision

    Modalities: Text, Image1M context

    $6.00/1M in

    $30.00/1M out

  • Claude Opus 4.8

    AnthropicText & Chat

    New

    The most capable model of Anthropic's Opus 4 series. State of the art on long-horizon agentic work, coding and knowledge tasks, with a 1M-token context window at standard pricing.

    claudeopusagentic

    Modalities: Text, Image1M context

    $6.00/1M in

    $30.00/1M out

  • Claude Opus 5.5

    AnthropicText & Chat

    New

    Anthropic's current Opus model for long-running agentic coding and knowledge work. 1M-token context window, up to 128K output tokens, adaptive thinking that is always on.

    claudeopusagentic

    Modalities: Text, Image1M context

    $4.80/1M in

    $24.00/1M out

  • Claude Sonnet 4.6

    AnthropicMultimodal

    New

    Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    balancedproductionagentic

    Modalities: Text, Image1M context

    $3.60/1M in

    $18.00/1M out

  • Claude Sonnet 5

    AnthropicText & Chat

    New

    Anthropic's Sonnet model with the best combination of speed and intelligence. 1M-token context window, up to 128K output tokens, adaptive thinking.

    claudesonnetbalanced

    Modalities: Text, Image1M context

    $2.40/1M in

    $12.00/1M out

  • DeepSeek V4 Pro

    DeepSeekText & Chat

    New

    DeepSeek's April 2026 flagship. 1.6T MoE / 49B active params, 1M context, rivals top closed-source models on STEM and coding at a fraction of the price.

    open-weightsmoecoding

    Modalities: Text1M context

    $1.584/1M in

    $4.752/1M out

  • Gemini 2.5 Pro

    Google DeepMindText & Chat

    New

    Google's latest thinking model. Excels at reasoning, coding, math, and science with massive context window.

    reasoningcodingmultimodal

    Modalities: Text1M context

    $1.50/1M in

    $12.00/1M out

  • Gemini 3 Flash

    Google DeepMindMultimodal

    New

    Google's April 2026 fast multimodal model. Combines Gemini 3 Pro's reasoning with Flash-tier latency and price. Default model in the Gemini app.

    balancedmultimodallow-latency

    Modalities: Text, Image, Video, Audio1M context

    $0.60/1M in

    $3.60/1M out

  • Gemini 3.1 Pro

    Google DeepMindMultimodal

    New

    Google DeepMind's February 2026 flagship. 1M-token (1,048,576) context, native multimodal (text/image/audio/video), Deep Think reasoning.

    multimodaldeep-thinklong-context

    Modalities: Text, Image, Video, Audio1M context

    $2.40/1M in

    $14.40/1M out

  • Google Veo 3.1

    Google DeepMindVideo

    New

    Latest Veo with image-to-video and context-aware audio

    audioi2v

    Modalities: Video

    $0.48/s

  • GPT-4.1

    OpenAIText & Chat

    New

    OpenAI's newest flagship model. Improved reasoning, instruction following, and coding over GPT-4o.

    codingreasoning

    Modalities: Text1M context

    $2.40/1M in

    $9.60/1M out

  • GPT-5.4

    OpenAIMultimodal

    New

    OpenAI's unified flagship combining GPT and o-series reasoning into one model. 1M context, multimodal, top SWE-Bench Pro and OSWorld scores.

    reasoningagenticvision

    Modalities: Text, Image1.1M context

    $3.00/1M in

    $18.00/1M out

  • GPT-5.4 Mini

    OpenAIMultimodal

    New

    OpenAI's efficient mid-tier model. 2x faster than its predecessor, 400k context, approaches GPT-5.4 quality on SWE-Bench Pro at a fraction of the cost.

    balancedcost-efficientvision

    Modalities: Text, Image400K context

    $0.90/1M in

    $5.40/1M out

  • GPT-5.5

    OpenAIText & Chat

    New

    OpenAI's current flagship chat model (released April 2026). Strongest general reasoning, coding and tool use in the GPT-5 line, with vision input and a large context window.

    gpt-5reasoningvision

    Modalities: Text, Image400K context

    $6.00/1M in

    $36.00/1M out

  • Kling v3

    Kuaishou (Kling)Video

    New

    Cinematic video up to 15s with multi-shot and native audio

    audioi2v

    Modalities: Video

    $0.2688/s

  • Kling v3 Omni

    Kuaishou (Kling)Video

    New

    Most versatile: multi-reference images, video editing, native audio

    audioi2vediting

    Modalities: Video

    $0.2688/s

  • o3-mini

    OpenAIText & Chat

    New

    OpenAI's reasoning model optimized for STEM tasks, coding, and math. Uses chain-of-thought reasoning.

    reasoningcodingmath

    Modalities: Text200K context

    $1.32/1M in

    $5.28/1M out

  • Runway Gen 4.5

    RunwayVideo

    New

    Top-ranked for motion quality and visual fidelity

    top-quality

    Modalities: Video

    $0.144/s

  • BLIP

    SalesforceMultimodal

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    blipcaptioningvqa

    Modalities: Text, Image

    ≈ $0.00030/run

  • CLIP Interrogator

    CommunityMultimodal

    pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    clip-interrogatorcaptioningtagging

    Modalities: Text, Image

    ≈ $0.0457/run

  • Depth Anything v2

    CommunityMultimodal

    Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    depthvision-understandingopen-weights

    Modalities: Text, Image

    ≈ $0.0050/run

  • FLUX 1.1 Pro

    Black Forest LabsImage

    Black Forest Labs' flagship text-to-image model. Faster generation than FLUX.1 Pro at higher prompt adherence, with strong photorealism and reliable spatial composition. Runs as a hosted Replicate model.

    fluxtext-to-imagephotorealism

    Modalities: Text, Image

    $0.048/image

  • Flux 1.1 Pro Ultra

    Black Forest LabsImage

    FLUX 1.1 Pro in ultra mode. Up to 4 megapixel images with raw mode for photorealism.

    high-qualityphotorealistic

    Modalities: Image

    $0.072/image

  • Flux Dev

    Black Forest LabsImage

    Black Forest Labs' development model. Fast, high-quality image generation with LoRA support.

    fastlora

    Modalities: Image

    $0.0384/image

  • Google Imagen 4

    Google DeepMindImage

    Google DeepMind's Imagen 4 text-to-image model, hosted on Replicate. Sharp detail, accurate text rendering, and strong prompt adherence across photographic and illustrated styles. Outputs up to 2K resolution.

    imagentext-to-imagephotorealism

    Modalities: Text, Image

    $0.048/image

  • Google Veo 2

    Google DeepMindVideo

    Google's state-of-the-art video generation model. Simulates real-world physics with various visual styles.

    high-quality

    Modalities: Video

    $0.60/s

  • Google Veo 3 (Replicate)

    Google DeepMindVideo

    Google's Veo 3 served via Replicate. Text-to-video with native synchronized audio generation. High-fidelity motion and scene coherence in short clips.

    veotext-to-videoaudio

    Modalities: Text, Image, Video

    $0.48/s

  • GPT-4o

    OpenAIText & Chat

    OpenAI's most capable multimodal model. Excellent for complex reasoning, coding, and creative tasks.

    fastmultimodal

    Modalities: Text128K context

    $3.00/1M in

    $12.00/1M out

  • HunyuanVideo

    TencentVideo

    Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.

    hunyuanvideotext-to-video

    Modalities: Text, Video

    ≈ $3.06/run

  • Icons (SDXL Flat Pop)

    CommunityImage

    SDXL fine-tune by galleri5 for slick flat icons and pop constructivist graphics with thick edges. Trained on Bing generations, it produces clean single-subject icon art that suits app icons, badges and UI glyphs. Raster output, not true vector.

    iconlogosdxl

    Modalities: Image

    ≈ $0.0094/run

  • Ideogram v3 Quality

    IdeogramImage

    The highest-quality tier of Ideogram v3. Improved photorealism and prompt adherence over v2 while keeping Ideogram's best-in-class text rendering. Supports style references and inline text layout.

    text-to-imagetypographyphotorealism

    Modalities: Text, Image

    $0.108/image

  • Incredibly Fast Whisper

    CommunitySpeech-to-Text

    Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    whisperstttranscription

    Modalities: Text, Audio

    ≈ $0.0056/run

  • InstantID

    CommunityImage

    InstantID makes realistic portraits of a real person from a single reference photo without per-user training. Combines a face encoder with an IdentityNet adapter on SDXL to keep identity and pose while following a text prompt, so it is fast and tuning-free.

    avatarportraitinstant-id

    Modalities: Text, Image

    ≈ $0.084/run

  • Kling v2.1

    Kuaishou (Kling)Video

    Kuaishou's Kling v2.1, generating 5 and 10 second videos at 720p or 1080p from text or an image. Known for cinematic camera work and realistic physical motion, available on Replicate via the official KwaiVGI account.

    videotext-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.060/s

  • Kling v2.1 Master

    Kuaishou (Kling)Video

    Kuaishou's premium Kling v2.1 Master. Generates 1080p 5s and 10s clips from text or an image with strong dynamics and prompt adherence. The top tier of the Kling 2.1 family.

    text-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.336/s

  • MiniMax Hailuo 02

    MiniMaxVideo

    MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.

    hailuotext-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.324/video

  • MusicGen

    MetaAudio & Music

    Meta's music generation model. Generate up to 1 minute of music from text descriptions.

    music

    Modalities: Audio

    ≈ $0.0529/run

  • OpenAI's highest-quality embedding model. Returns 3072-dim vectors by default and supports reducing dimensions via the dimensions parameter. Outperforms text-embedding-3-small and the older ada-002 on MTEB and multilingual MIRACL retrieval benchmarks, for cases where accuracy matters more than cost.

    embeddingretrievalrag

    Modalities: Text, Vectors8.2K context

    $0.156/1M in

  • OpenAI's small, low-cost embedding model. Returns 1536-dim vectors by default and supports shortening output dimensions via the dimensions parameter without retraining. Replaced text-embedding-ada-002 with better retrieval quality at a fraction of the price, and is the default choice for general-purpose semantic search and RAG.

    embeddingretrievalrag

    Modalities: Text, Vectors8.2K context

    $0.024/1M in

  • Turns any single selfie into a clean professional headshot using FLUX Kontext image editing. Keeps the person's face while swapping to business attire, a studio background and even lighting. Aimed at LinkedIn-style profile photos.

    avatarportraitheadshot

    Modalities: Image

    $0.048/image

  • Recraft 20B SVG

    RecraftImage

    Recraft's faster, cheaper vector model. Outputs editable SVG paths instead of raster pixels, so logos, icons and flat illustrations scale to any size without blur. Defaults to a vector_illustration style and supports line art and engraving looks. Hosted API only.

    svgvectorlogo

    Modalities: Image

    $0.0528/image

  • Recraft Vectorize

    RecraftImage

    Recraft's raster-to-vector converter. Takes a PNG or JPG and traces it into a clean SVG with precise vector paths, aimed at logos, icons and graphics that need to scale. Image-to-SVG counterpart to Recraft's text-to-SVG models.

    svgvectorvectorize

    Modalities: Image

    $0.012/image

  • Runway Gen-4 Turbo

    RunwayVideo

    Runway's Gen-4 Turbo on Replicate. Fast image-to-video generation producing 5s and 10s clips at 720p with strong character and scene consistency across shots.

    gen-4image-to-videofast

    Modalities: Text, Image, Video

    $0.060/s

  • Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.

    segmentationvision-understandingopen-weights

    Modalities: Text, Image, Video

    ≈ $0.018/run

  • Sticker Maker

    CommunityImage

    fofr's sticker generator that outputs graphics with transparent backgrounds, so the result drops straight into chat apps or print sheets. Runs an SDXL-based pipeline at high speed (default 17 steps) and returns die-cut style art without manual background removal.

    stickertransparentfofr

    Modalities: Image

    ≈ $0.0055/run

  • Whisper

    OpenAISpeech-to-Text

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    whisperstttranscription

    Modalities: Text, Audio

    ≈ $0.0034/run

  • Claude Haiku 4.5

    AnthropicMultimodal

    New

    Anthropic's fastest and cheapest 4.x model. Strong vision and tool use at ultra-low latency, ideal for high-concurrency workloads.

    cost-efficientlow-latencyvision

    Modalities: Text, Image200K context

    $1.20/1M in

    $6.00/1M out

  • DeepSeek V4 Flash

    DeepSeekText & Chat

    New

    Efficiency-optimized variant of DeepSeek V4. 284B MoE / 13B active, 1M context, ultra-low pricing for high-throughput workloads.

    open-weightsmoecost-efficient

    Modalities: Text1M context

    $0.36/1M in

    $1.44/1M out

  • DeepSeek V4.1 Flash

    DeepSeekText & Chat

    New

    DeepSeek's current Flash model (API name deepseek-flash, model version DeepSeek-V4.1-Flash): 1M-token context, up to 384K output tokens, JSON output, tool calls and vision input.

    open-weightscost-efficientlong-context

    Modalities: Text, Image1M context

    $0.36/1M in

    $1.44/1M out

  • Fibo

    Bria AIImage

    New

    SOTA Open source model trained on licensed data, transforming intent into structured control for precise, high-quality AI image generation in enterprise and agentic workflows.

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • Flux Fast

    CommunityImage

    New

    This is the fastest Flux endpoint in the world.

    prunaaitext-to-image

    Modalities: Text, Image

    $0.0060/image

  • Flux Kontext Max

    Black Forest LabsImage

    New

    A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural language prompts

    text-to-image

    Modalities: Text, Image

    $0.096/image

  • Flux Kontext Pro

    Black Forest LabsImage

    New

    A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming images through natural language

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • Flux Krea Dev

    Black Forest LabsImage

    New

    An opinionated text-to-image model from Black Forest Labs in collaboration with Krea that excels in photorealism. Creates images that avoid the oversaturated "AI look".

    text-to-image

    Modalities: Text, Image

    $0.030/image

  • Flux Pro

    Black Forest LabsImage

    New

    State-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity.

    text-to-image

    Modalities: Text, Image

    $0.066/image

  • Gen4 Image

    RunwayImage

    New

    Runway's Gen-4 Image model with references. Use up to 3 reference images to create the exact image you need. Capture every angle.

    runwaymltext-to-image

    Modalities: Text, Image

    $0.096/image

  • Google Veo 3 Fast

    Google DeepMindVideo

    New

    Faster cheaper Veo 3 with audio

    fastaudio

    Modalities: Video

    $0.18/s

  • Google Veo 3.1 Fast

    Google DeepMindVideo

    New

    Faster Veo 3.1 with image-to-video and audio

    fastaudioi2v

    Modalities: Video

    $0.18/s

  • Gpt Image 1.5

    OpenAIImage

    New

    OpenAI's latest image generation model with better instruction following and adherence to prompts

    text-to-image

    Modalities: Text, Image

    $0.1632/image

More models (213)

211 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

All AI models, one API — your unified catalog

Railwail gives you a single OpenAI-compatible endpoint for AI models from many providers — GPT-5.5 and Claude Opus 4.8 for reasoning, Gemini 3.1 Pro for long contexts, FLUX 1.1 Pro for photorealistic images, Veo 3.1 for video with synced audio, Whisper for speech-to-text, OpenAI TTS for voice and OpenAI text-embedding-3 for embeddings. You pick a model, change one parameter in your request, and ship. No new SDK, no new auth flow, no provider lock-in — the catalog above lists every model with its current price in USD and whether it can currently be run.

The pricing is transparent and on-demand: you see the price before you call, you pay per token (or per image, or per second for video and audio), and there are no monthly minimums, no seat fees, and no surprise overage charges. Accounts that sign up with Google get trial credits for the cheaper models, so you can try a few prompts before topping up. Switching between flagships is a one-line change: replace `model: "gpt-5-4"` with `model: "claude-sonnet-4-6"` or `model: "gemini-3-1-pro"` and the rest of your code keeps working. That same surface covers cheap, fast budget tiers like GPT-5.4 Nano, Claude Haiku, Gemini 3 Flash and DeepSeek V4 Flash when latency or unit cost matters more than peak quality.

The platform runs on servers in Germany, Railwail does not train on customer prompts, and every model page names the provider that processes the request, so your compliance team can check each one. Compared with OpenRouter or Together AI, the difference is an operator based in Germany, a prepaid balance in USD and one catalog that covers ten categories — text, image, video, audio, text-to-speech, speech-to-text, embeddings, code, multimodal, and vision-language-action robotics — so a single integration handles your chatbot, your image pipeline, your transcripts, and your RAG retriever without juggling five SDKs.

Typical tasks

  • Customer support chatbots
  • Code generation in IDEs
  • Long-document summarization
  • Image generation pipelines
  • Multilingual translation
  • Voice transcription at scale

Model comparisons

Frequently asked questions

How do I choose the right AI model for my project?

Start from your workload, not from the leaderboard. For high-volume classification or extraction, pick a fast budget tier (Gemini Flash, GPT-5 Mini, Claude Haiku, DeepSeek V3) — quality is 90% of flagship at 5-10% of the cost. For user-facing reasoning, long-form writing, or code review, pick a flagship (GPT-5, Claude 4.6 Sonnet, Gemini 3 Pro). For images use FLUX 1.1 Pro or Imagen 4; for video Veo 3.1 or Runway Gen-4; for transcripts Whisper or WhisperX. The catalog above can be filtered by category and sorted by price, so you can match the trade-off triangle that matters to your product.

What's the difference between LLMs, image models, and VLA models?

LLMs (large language models) read and write text — GPT-5, Claude 4.6, Gemini 3. Image models like FLUX, Imagen, and Stable Diffusion turn a text prompt into a raster image. Video models like Veo 3 and Kling produce motion clips. Audio models cover music (MusicGen, Stable Audio Open) and sound effects. Speech-to-text turns audio into transcripts; text-to-speech goes the other way. Embedding models compress meaning into vectors for retrieval and search. Code models are LLMs fine-tuned for programming. Multimodal models accept text plus images or video as input. VLA (vision-language-action) models close the loop with motor control — they output robot actions, not text. Railwail lists all of them in one catalog; text, image, video, speech and embedding models run through the same OpenAI-compatible API.

How does Railwail's pricing work?

Pure pay-as-you-go in USD from a prepaid balance (1 credit = 0.01 USD). Text models bill per input and output token at rates shown on every model page; a typical prompt costs fractions of a cent. Image generation bills per image, video per second of output or per clip, speech per 1,000 characters, and some open-weights models by GPU time — the model page lists the exact rate before you call. No monthly minimums, no seat fees, no commitment. The dashboard shows your usage and transaction history, and in the settings you can set a monthly spend limit; new runs are blocked once it is reached.

Do I get free credits to start?

Accounts that sign up with Google receive 10 trial credits (0.10 USD). They can be used 24 hours after sign-up on runs that cost up to 2 credits each, which covers short prompts on the cheaper text models and budget image models such as FLUX Schnell. Accounts created by email or GitHub start without trial credits. No credit card is required to sign up and there is no auto-conversion to paid: when your balance runs out, calls simply stop until you top up (from 5 USD).

How does Railwail compare to OpenRouter or Together AI?

The OpenAI-compatible API surface is similar, so migrating from OpenRouter is mostly a base-URL change. Railwail is operated from Germany, runs its platform on servers in Germany and bills a prepaid balance in USD, with VAT handled at billing time and no monthly minimums or seat fees. Like the other gateways, it forwards each request to the provider of the chosen model, and that provider may process the data outside the EU. Features such as streaming, tool calling and automatic provider failover are not available through Railwail yet — check that before you migrate a workload that depends on them.

Is Railwail GDPR/DSGVO compliant?

Railwail runs its platform, database and request logs on servers in Germany and does not use customer prompts for model training. Each request is processed by the provider of the chosen model; many providers are located outside the EU, and every model page names the provider so your DPO can check it before approving a model. OpenAI and Anthropic state that they do not train on API data by default; other providers' terms differ. Business customers can ask for a data processing agreement, and requests for access or deletion can be sent to [email protected]. Details are on the trust page and in the privacy policy.

Can I switch between models without code changes?

Yes — the whole point of the unified catalog. Railwail exposes an OpenAI-compatible endpoint, so SDKs that talk to OpenAI (openai-python, openai-node, the official SDKs in other languages, plus LangChain, LlamaIndex, Vercel AI SDK and similar) work by changing just the `baseURL` and the API key. After that, swapping `model: "gpt-5-4"` for `model: "claude-sonnet-4-6"` or `model: "gemini-3-1-pro"` is a one-line change; messages, temperature, max_tokens, top_p and stop keep the same shape. `tools`, `response_format` and `stream` are currently not forwarded, and provider-specific extras such as Anthropic's prompt-caching headers are not passed through.

What happens if a provider has downtime?

Each model in the catalog is served by one named upstream provider. If that provider fails, the request returns an error and the credits reserved for it are refunded to your balance. Automatic failover to other providers is not offered as a feature, so build retries or a second model into your own code if a workload must not fail. There is no public status page at the moment.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.