Replicate

San Francisco, USAFounded 2019
241 models

Replicate is an open-model hosting platform that serves thousands of open-source models including Flux, Stable Diffusion, Llama, and Whisper variants via a unified API.

241 models from Replicate on Railwail

Access every Replicate model through Railwail's OpenAI-compatible API.

241 models

  • BLIP

    MultimodalSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    $0.00030 / run
    replicateblipcaptioning
  • CLIP Interrogator

    MultimodalCommunity

    pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    $0.046 / run
    replicateclip-interrogatorcaptioning
  • Depth Anything v2

    MultimodalCommunity

    Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    $0.0050 / run
    replicatedepthvision-understanding
  • FLUX 1.1 Pro

    ImageBlack Forest Labs

    Black Forest Labs' flagship text-to-image model. Faster generation than FLUX.1 Pro at higher prompt adherence, with strong photorealism and reliable spatial composition. Runs as a hosted Replicate model.

    $0.048 / run
    replicatefluxblack-forest-labs
  • Flux 1.1 Pro Ultra

    ImageBlack Forest Labs

    FLUX 1.1 Pro in ultra mode. Up to 4 megapixel images with raw mode for photorealism.

    $0.072 / run
    high-qualityphotorealistic
  • Flux Dev

    ImageBlack Forest Labs

    Black Forest Labs' development model. Fast, high-quality image generation with LoRA support.

    $0.038 / run
    popularfastlora
  • Google Imagen 4

    ImageGoogle DeepMind

    Google DeepMind's Imagen 4 text-to-image model, hosted on Replicate. Sharp detail, accurate text rendering, and strong prompt adherence across photographic and illustrated styles. Outputs up to 2K resolution.

    $0.048 / run
    replicategoogleimagen
  • Google Veo 2

    VideoGoogle DeepMind

    Google's state-of-the-art video generation model. Simulates real-world physics with various visual styles.

    $4.80 / run
    high-qualitypopular
  • Google Veo 3 (Replicate)

    VideoGoogle DeepMind

    Google's Veo 3 served via Replicate. Text-to-video with native synchronized audio generation. High-fidelity motion and scene coherence in short clips.

    $3.84 / run
    replicategoogleveo
  • Google Veo 3.1

    VideoGoogle DeepMind
    New

    Latest Veo with image-to-video and context-aware audio

    $1.92 / run
    popularaudioi2v
  • HunyuanVideo

    VideoTencent

    Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.

    $3.06 / run
    replicatetencenthunyuan
  • Icons (SDXL Flat Pop)

    ImageCommunity

    SDXL fine-tune by galleri5 for slick flat icons and pop constructivist graphics with thick edges. Trained on Bing generations, it produces clean single-subject icon art that suits app icons, badges and UI glyphs. Raster output, not true vector.

    $0.0094 / run
    replicateiconlogo
  • Ideogram v3 Quality

    ImageIdeogram

    The highest-quality tier of Ideogram v3. Improved photorealism and prompt adherence over v2 while keeping Ideogram's best-in-class text rendering. Supports style references and inline text layout.

    $0.11 / run
    replicateideogramtext-to-image
  • Incredibly Fast Whisper

    Speech-to-TextCommunity

    Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    $0.0056 / run
    replicatewhisperstt
  • InstantID

    ImageCommunity

    InstantID makes realistic portraits of a real person from a single reference photo without per-user training. Combines a face encoder with an IdentityNet adapter on SDXL to keep identity and pose while following a text prompt, so it is fast and tuning-free.

    $0.084 / run
    avatarportraitinstant-id
  • Kling v2.1

    VideoKuaishou (Kling)

    Kuaishou's Kling v2.1, generating 5 and 10 second videos at 720p or 1080p from text or an image. Known for cinematic camera work and realistic physical motion, available on Replicate via the official KwaiVGI account.

    $0.30 / run
    replicatekuaishoukling
  • Kling v2.1 Master

    VideoKuaishou (Kling)

    Kuaishou's premium Kling v2.1 Master. Generates 1080p 5s and 10s clips from text or an image with strong dynamics and prompt adherence. The top tier of the Kling 2.1 family.

    $1.68 / run
    replicatekuaishoukling
  • Kling v3

    VideoKuaishou (Kling)
    New

    Cinematic video up to 15s with multi-shot and native audio

    $1.34 / run
    popularaudioi2v
  • Kling v3 Omni

    VideoKuaishou (Kling)
    New

    Most versatile: multi-reference images, video editing, native audio

    $1.34 / run
    popularaudioi2v
  • MiniMax Hailuo 02

    VideoMiniMax

    MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.

    $0.32 / run
    replicateminimaxhailuo
  • MusicGen

    Audio & MusicMeta

    Meta's music generation model. Generate up to 1 minute of music from text descriptions.

    $0.053 / run
    musicpopular
  • Turns any single selfie into a clean professional headshot using FLUX Kontext image editing. Keeps the person's face while swapping to business attire, a studio background and even lighting. Aimed at LinkedIn-style profile photos.

    $0.048 / run
    avatarportraitheadshot
  • Recraft 20B SVG

    ImageRecraft

    Recraft's faster, cheaper vector model. Outputs editable SVG paths instead of raster pixels, so logos, icons and flat illustrations scale to any size without blur. Defaults to a vector_illustration style and supports line art and engraving looks. Hosted API only.

    $0.053 / run
    replicaterecraftsvg
  • Recraft Vectorize

    ImageRecraft

    Recraft's raster-to-vector converter. Takes a PNG or JPG and traces it into a clean SVG with precise vector paths, aimed at logos, icons and graphics that need to scale. Image-to-SVG counterpart to Recraft's text-to-SVG models.

    $0.012 / run
    svgvectorrecraft
  • Runway Gen 4.5

    VideoRunway
    New

    Top-ranked for motion quality and visual fidelity

    $0.72 / run
    populartop-quality
  • Runway's Gen-4 Turbo on Replicate. Fast image-to-video generation producing 5s and 10s clips at 720p with strong character and scene consistency across shots.

    $0.30 / run
    replicaterunwaygen-4
  • Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.

    $0.018 / run
    replicatesegmentationmeta
  • Sticker Maker

    ImageCommunity

    fofr's sticker generator that outputs graphics with transparent backgrounds, so the result drops straight into chat apps or print sheets. Runs an SDXL-based pipeline at high speed (default 17 steps) and returns die-cut style art without manual background removal.

    $0.0055 / run
    replicatestickertransparent
  • Whisper

    Speech-to-TextOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    $0.0034 / run
    replicateopenaiwhisper
  • Background removal model from 851-Labs that outputs a clean cutout with a transparent alpha channel. One of the most-run background removers on Replicate, handles people, products and objects on busy backgrounds.

    $0.00050 / run
    background-removal851-labscutout
  • Product advertising photo generator. You upload a cut-out product shot and a prompt describing the scene; it places the product on a new generated background with matching lighting and shadows, so a plain packshot becomes an ecommerce or ad-ready hero image without a photo studio.

    $0.067 / run
    productecommerceproduct-photo
  • AnimateDiff

    VideoCommunity

    Plug-and-play motion module that animates personalized Stable Diffusion models without further training. 16-frame clips at 512x512.

    $0.12 / run
    replicateanimationanimatediff
  • AnimateDiff Lightning

    VideoCommunity

    ByteDance distillation of AnimateDiff. 4-step sampling for over 10x faster inference at comparable quality to multi-step base model.

    $0.070 / run
    replicateanimationbytedance
  • AudioLDM 2

    Text-to-SpeechHaohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    $0.016 / run
    audioldmmusic-generationdiffusion
  • AuraFlow v0.3

    ImageCommunity

    fal.ai's fully open-source 6.8B flow-based text-to-image model. Up to 1536x1536 resolution.

    $0.065 / run
    auraflowtext-to-imageopen-weights
  • BiRefNet high-resolution dichotomous image segmentation for background removal. Bilateral reference network that produces sharp matting on fine detail like hair, fur and thin structures, often cleaner than older U2Net or rembg models.

    $0.0036 / run
    background-removalbirefnetsegmentation
  • BRIA AI's commercial background removal model trained on fully licensed data. Produces accurate cutouts for e-commerce and design, with attention to clean edges around products and people.

    $0.022 / run
    background-removalbriaecommerce
  • Microsoft Research pipeline by Ziyu Wan et al. that restores scanned old photos, removing scratches, dust and fading and optionally enhancing faces in one pass.

    $0.010 / run
    restoreold-photoscratch-removal
  • Cartoonify

    ImageCommunity

    catacolabs Cartoonify turns a photo into a flat cartoon illustration. Takes a single image and returns a stylized cartoon version with clean shapes and bold outlines. Straightforward one-input model for avatars and profile pictures.

    $0.096 / run
    avatarportraitcartoon
  • Content-Consistent Super-Resolution model. Reduces hallucination compared to typical diffusion-based upscalers while keeping perceptual quality high.

    $0.28 / run
    replicateupscalingimage-restore
  • Chatterbox

    Text-to-SpeechResemble AI

    Resemble AI's open Chatterbox TTS. Zero-shot voice cloning from a short audio prompt with an exaggeration control for emotion intensity, plus CFG weight to balance pacing and fidelity.

    $0.030 / run
    replicateresemble-aitts
  • Clarity Upscaler

    ImageCommunity

    High-resolution image upscaler with creative detail re-imagination via SD-based hallucination. Strong for photography and product shots.

    $0.024 / run
    replicateupscalingcreative
  • Meta's 13B Code Llama tuned for instruction following. A faster mid-size option for code generation and completion, supporting infilling for inserting code at a cursor position. Served on Replicate per call.

    $0.0065 / run
    metacode-llamacoding
  • Meta's 34B Code Llama tuned for instruction following. A balance of size and quality for code generation, completion, and explanation, with strong coverage of Python, JavaScript, and other common languages. Runs on Replicate per call.

    $0.041 / run
    metacode-llamacoding
  • Meta's largest Code Llama, a 70B Llama-2 derivative specialized for programming and tuned to follow instructions in chat form. Handles code generation, completion, and explanation across common languages. Served on Replicate as a per-call endpoint.

    $0.050 / run
    metacode-llamacoding
  • Meta's smallest Code Llama at 7B parameters, tuned for instruction following. The cheapest and fastest member of the family for quick code generation, completion, and infilling. Served on Replicate per call.

    $0.0074 / run
    metacode-llamacoding
  • CodeFormer

    ImageCommunity

    Robust face-restoration model using a transformer-based codebook prior. Handles severe degradation, occlusion, and old-photo restoration with adjustable fidelity-quality tradeoff.

    $0.0041 / run
    replicateface-restoreupscaling
  • CogVideoX-5B

    VideoCommunity

    CogVideoX-5B from Tsinghua/Zhipu AI, an open 5B-parameter text-to-video diffusion transformer. Generates 6-second 720p clips with coherent motion and is widely used in research for its open weights and reproducibility.

    $0.60 / run
    replicatecogvideoxzhipu
  • CogVideoX-5B (open)

    VideoZhipu AI

    Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

    $0.41 / run
    zhiputsinghuacogvideox
  • CogVLM2 19B

    MultimodalCommunity

    Tsinghua CogVLM2 19B with Llama-3 8B base plus 11B vision expert. Strong document understanding and visual reasoning, 8k context.

    $0.011 / run
    replicatemultimodalvision-understanding
  • Consistent Character

    ImageCommunity

    fofr's model generates the same character in many poses and angles from one reference image. Useful for building an avatar set or character sheet where the face and design stay consistent across outputs. Can produce a grid or individual images.

    $0.0038 / run
    avatarportraitcharacter
  • ControlNet Canny

    ImageCommunity

    ControlNet conditioned on Canny edge maps. Preserves composition and outlines while restyling with Stable Diffusion 1.5 or SDXL backbones.

    $0.070 / run
    replicatestyle-transferimage-edit
  • ControlNet Depth

    ImageCommunity

    ControlNet conditioned on depth maps. Preserves the 3D scene layout while letting the prompt change style, lighting and content.

    $0.014 / run
    replicatestyle-transferimage-edit
  • DDColor

    ImageCommunity

    DDColor by Xiaoyang Kang et al. colorizes black-and-white photos using dual decoders that jointly learn pixel colors and semantic color queries, giving vivid and natural results on old images.

    $0.0012 / run
    colorizerestoreddcolor
  • Quantized GGUF build of DeepSeek's 33B code model, trained on roughly 2T tokens that are about 87 percent code. Designed for repository-level completion and project-aware generation thanks to a 16k context window. Runs on Replicate as a per-call endpoint.

    $0.0030 / run
    deepseekcodinginstruct
  • DeepSeek-VL 7B

    MultimodalDeepSeek

    DeepSeek-VL 7B chat model. Vision-language model with hybrid vision encoder and strong real-world visual question answering performance.

    $0.0086 / run
    replicatemultimodalvision-understanding
  • Donut Document

    MultimodalCommunity

    Naver CLOVA Donut OCR-free document-understanding transformer. End-to-end JSON extraction from forms, receipts and invoices without explicit OCR.

    $0.065 / run
    replicateocrvision-understanding
  • Dots OCR

    MultimodalCommunity

    Rednote Hilab Dots OCR. End-to-end document parsing model with layout, text and reading-order prediction in one transformer.

    $0.013 / run
    replicateocrvision-understanding
  • DreamGaussian

    ImageCommunity

    Generative Gaussian-splatting model for fast image-to-3D synthesis. Produces textured meshes in two minutes via differentiable rasterization.

    $0.10 / run
    replicate3d-generationimage-to-3d
  • EasyOCR

    MultimodalCommunity

    JaidedAI EasyOCR. Simple Python OCR wrapper supporting 80+ languages with deep-learning text detection and recognition.

    $0.0014 / run
    replicateocrvision-understanding
  • EchoMimic

    VideoCommunity

    Ant Group EchoMimic. Lifelike audio-driven portrait animation with editable landmark conditioning for fine-grained motion control.

    $0.48 / run
    replicatelipsyncant-group
  • Face to Many

    ImageCommunity

    fofr's face stylizer converts a face photo into 3D render, emoji, pixel art, video-game character, claymation or toy styles. Uses InstantID plus style LoRAs on SDXL to keep the likeness while applying a chosen art style. Popular for fun avatars.

    $0.011 / run
    avatarportraitstylize
  • Face to Sticker

    ImageCommunity

    fofr's model turns a face photo into a die-cut sticker with a white border and transparent background. Uses InstantID to hold the likeness and outputs a clean PNG suitable for chat stickers or print. Simple single-image input.

    $0.029 / run
    avatarportraitsticker
  • Fibo

    ImageBria AI
    New

    SOTA Open source model trained on licensed data, transforming intent into structured control for precise, high-quality AI image generation in enterprise and agentic workflows.

    $0.048 / run
    replicatebriatext-to-image
  • FILM Frame Interpolation

    VideoGoogle Research

    Google FILM frame interpolation. Synthesizes high-quality intermediate frames between near-duplicate inputs, designed for large motion gaps.

    $0.0016 / run
    replicateupscaleframe-interpolation
  • Florence-2 Large

    MultimodalCommunity

    Microsoft Florence-2 Large. Unified prompt-based vision foundation model for captioning, detection, segmentation and OCR with a single 770M-param backbone.

    $0.0012 / run
    replicatemultimodalvision-understanding
  • Flux Fast

    ImageCommunity
    New

    This is the fastest Flux endpoint in the world.

    $0.0060 / run
    replicateprunaaitext-to-image
  • Flux Kontext Max

    ImageBlack Forest Labs
    New

    A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural language prompts

    $0.096 / run
    replicateblack-forest-labstext-to-image
  • Flux Kontext Pro

    ImageBlack Forest Labs
    New

    A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming images through natural language

    $0.048 / run
    replicateblack-forest-labstext-to-image
  • Flux Krea Dev

    ImageBlack Forest Labs
    New

    An opinionated text-to-image model from Black Forest Labs in collaboration with Krea that excels in photorealism. Creates images that avoid the oversaturated "AI look".

    $0.030 / run
    replicateblack-forest-labstext-to-image
  • Flux Pro

    ImageBlack Forest Labs
    New

    State-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity.

    $0.066 / run
    replicateblack-forest-labstext-to-image
  • FLUX PuLID

    ImageCommunity

    PuLID identity customization running on FLUX.1-dev. Inserts a face from one reference photo into prompt-driven scenes using contrastive alignment, giving higher likeness and detail than SDXL-era ID adapters. Good for realistic avatars and character portraits.

    $0.018 / run
    avatarportraitpulid
  • FLUX.1 [dev]

    ImageBlack Forest Labs

    The open-weight 12B rectified-flow transformer from Black Forest Labs. Close to FLUX Pro quality with a guidance-distilled checkpoint released under a non-commercial license. The most widely fine-tuned base in the FLUX family.

    $0.030 / run
    replicatefluxblack-forest-labs
  • FLUX.1 [schnell]

    ImageBlack Forest Labs

    The fastest FLUX model from Black Forest Labs, distilled to produce images in 1 to 4 steps. Apache 2.0 licensed for commercial use. Built for high-volume generation and real-time previews.

    $0.0036 / run
    replicatefluxblack-forest-labs
  • FLUX.1 Canny

    ImageBlack Forest Labs

    FLUX structural control via Canny edge maps. Preserve composition while restyling.

    $0.060 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1 Canny [dev]

    ImageBlack Forest Labs

    Open-weight edge-guided FLUX model from Black Forest Labs. Extracts Canny edges from a control image and regenerates it from your prompt while holding the original composition and outlines, so you can restyle a scene without changing its structure.

    $0.030 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1 Depth

    ImageBlack Forest Labs

    FLUX structural control via depth maps. Keep 3D scene layout while changing style/content.

    $0.060 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1 Depth [dev]

    ImageBlack Forest Labs

    Open-weight depth-guided FLUX model from Black Forest Labs. Derives a depth map from the control image and regenerates from your prompt while preserving 3D spatial layout, useful for re-texturing rooms, products, or scenes without moving objects.

    $0.030 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1 Fill

    ImageBlack Forest Labs

    Black Forest Labs' inpainting/outpainting model for FLUX. Fill masked regions with prompt-guided content.

    $0.060 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1 Fill [dev]

    ImageBlack Forest Labs

    Black Forest Labs' open-weight inpainting and outpainting model, guidance-distilled from FLUX.1 Fill [pro]. You supply an image plus a mask and a prompt; it fills the masked region or extends the canvas with prompt-guided content that matches lighting and texture.

    $0.048 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1 Kontext [dev]

    ImageBlack Forest Labs

    Open-weight version of FLUX.1 Kontext by Black Forest Labs. Instruction-based editing: pass an input image and a plain text edit ('change the jacket to red', 'remove the person on the left') and it applies the change while keeping the rest of the scene and identity consistent.

    $0.030 / run
    fluxkontextblack-forest-labs
  • FLUX.1 Redux

    ImageBlack Forest Labs

    FLUX image-variation adapter. Generate variations and remixes from a reference image.

    $0.0036 / run
    fluxblack-forest-labsimage-edit
  • FLUX.1-dev Inpainting

    ImageCommunity

    FLUX.1-dev inpainting wrapper that fills masked parts of an image from a prompt. Useful when you want FLUX-quality fills with a simple image plus mask plus prompt interface and adjustable mask strength.

    $0.090 / run
    fluximage-editinpainting
  • Gen4 Image

    ImageRunway
    New

    Runway's Gen-4 Image model with references. Use up to 3 reference images to create the exact image you need. Capture every angle.

    $0.096 / run
    replicaterunwaymltext-to-image
  • GFPGAN v1.4

    ImageTencent ARC

    Tencent ARC face-restoration GAN. Reconstructs realistic facial detail in low-quality or compressed photos using a pretrained StyleGAN2 prior.

    $0.0027 / run
    replicateface-restoreupscaling
  • GLPN Depth

    MultimodalCommunity

    Global-Local Path Networks depth-estimation model. Combines hierarchical transformer encoder with selective feature fusion for sharp boundaries.

    $0.0094 / run
    replicatedepthvision-understanding
  • Google Veo 3 Fast

    VideoGoogle DeepMind
    New

    Faster cheaper Veo 3 with audio

    $0.72 / run
    fastaudio
  • Google Veo 3.1 Fast

    VideoGoogle DeepMind
    New

    Faster Veo 3.1 with image-to-video and audio

    $0.72 / run
    fastaudioi2v
  • GOT-OCR 2.0

    MultimodalCommunity

    StepFun GOT-OCR 2.0. Unified end-to-end OCR-2.0 model handling text, formulas, charts, sheet music and geometric shapes in one architecture.

    $0.0012 / run
    replicateocrvision-understanding
  • Gpt Image 1.5

    ImageOpenAI
    New

    OpenAI's latest image generation model with better instruction following and adherence to prompts

    $0.16 / run
    replicateopenaitext-to-image
  • Gpt Image 2

    ImageOpenAI
    New

    OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.

    $0.15 / run
    replicateopenaitext-to-image
  • New

    OpenAI's fastest model for high-quality, everyday image generation. Generate and edit images from text and image inputs with strong instruction following and sharp text rendering.

    $0.30 / run
    replicateopenaitext-to-image
  • New

    OpenAI's most capable image model, built for workflows where editing precision matters most. Generate and edit images from text and image inputs with strong instruction following, sharp text rendering, and detailed control.

    $0.30 / run
    replicateopenaitext-to-image
  • IBM Granite 20B Code Instruct. Larger Granite code model balancing quality and inference cost for enterprise CI/CD code-review automation.

    $0.12 / 1M input
    replicatecode-generationibm
  • IBM Granite 8B Code Instruct. Trained on permissively-licensed code, strong on multi-language code completion and instruction-following.

    $0.060 / 1M input
    replicatecode-generationibm
  • New

    SOTA image model from xAI

    $0.024 / run
    replicatexaitext-to-image
  • New

    xAI's Grok Imagine Image 2.0 — text-to-image generation and editing with a quality control and output up to 2k

    $0.048 / run
    replicatexaitext-to-image
  • New

    xAI video with native audio and lip-sync, up to 15s

    $0.30 / run
    audioi2vxai
  • Grounded-SAM

    MultimodalCommunity

    Grounding DINO plus SAM. Open-vocabulary text-prompted detection and segmentation in one pipeline for fully-automatic mask generation.

    $0.0015 / run
    replicatesegmentationvision-understanding
  • Hailuo 2.3

    VideoMiniMax
    New

    Minimax model for realistic human motion and VFX

    $0.34 / run
    i2v1080p
  • Hidream L1 Fast

    ImageCommunity
    New

    This is an optimised version of the hidream-l1 model using the pruna ai optimisation toolkit!

    $0.0060 / run
    replicateprunaaitext-to-image
  • Html To Image

    ImageCommunity
    New

    Html To Image on Replicate (intelligent-utilities/html-to-image)

    $0.0012 / run
    replicateintelligent-utilitiestext-to-image
  • Hunyuan Image 2.1

    ImageTencent
    New

    Generate high-quality 2K resolution images from text prompts

    $0.024 / run
    replicatetencenttext-to-image
  • Hunyuan Image 3

    ImageTencent
    New

    A powerful native multimodal model for image generation (PrunaAI squeezed)

    $0.096 / run
    replicatetencenttext-to-image
  • Hunyuan3D 2.0

    ImageTencent

    Tencent's Hunyuan3D 2.0 image-to-3D pipeline. Two-stage shape and texture generation producing high-resolution textured meshes.

    $0.16 / run
    replicate3d-generationimage-to-3d
  • Hunyuan3D 2.1

    ImageCommunity
    New

    Refreshed Hunyuan3D 2.1 with improved texture fidelity and PBR-material support. Image-to-3D with textured GLB output.

    $0.26 / run
    replicate3d-generationimage-to-3d
  • Lvmin Zhang's IC-Light packaged by zsxkib. Relights a product or portrait from a text prompt or a chosen light direction while keeping the subject's shape and detail intact, so a flat product photo can be given studio, window, or dramatic side lighting without re-shooting.

    $0.078 / run
    productic-lightrelight
  • Idefics3 8B

    MultimodalCommunity

    Hugging Face Idefics3 8B. Llama-3 based open-source vision-language model with strong document QA and chart-understanding performance.

    $0.0012 / run
    replicatemultimodalvision-understanding
  • Ideogram v2

    ImageIdeogram

    Ideogram's text-to-image model known for accurate in-image text and typography. Handles posters, logos, and signage where other models garble lettering. Supports magic prompt expansion and multiple aspect ratios.

    $0.096 / run
    replicateideogramtext-to-image
  • Ideogram V2 Turbo

    ImageIdeogram
    New

    A fast image model with state of the art inpainting, prompt comprehension and text rendering.

    $0.060 / run
    replicateideogram-aitext-to-image
  • Ideogram V2A

    ImageIdeogram
    New

    Like Ideogram v2, but faster and cheaper

    $0.048 / run
    replicateideogram-aitext-to-image
  • Ideogram V2A Turbo

    ImageIdeogram
    New

    Like Ideogram v2 turbo, but now faster and cheaper

    $0.030 / run
    replicateideogram-aitext-to-image
  • New

    Balance speed, quality and cost. Ideogram v3 creates images with stunning realism, creative designs, and consistent styles

    $0.072 / run
    replicateideogram-aitext-to-image
  • Ideogram v3 Turbo

    ImageIdeogram

    Ideogram's fast v3 model, the fastest and cheapest tier of the v3 family. Known for accurate in-image text rendering and reliable typography, which most diffusion models still get wrong. Hosted API only.

    $0.036 / run
    replicateideogramtext-to-image
  • New

    Balance speed, quality and cost. Ideogram v4 creates images with stunning realism, creative designs, and consistent styles

    $0.072 / run
    replicateideogram-aitext-to-image
  • Ideogram V4 Quality

    ImageIdeogram
    New

    The highest quality Ideogram v4 model. v4 creates images with stunning realism, creative designs, and consistent styles

    $0.12 / run
    replicateideogram-aitext-to-image
  • IDM-VTON virtual try-on from the CVPR 2024 paper. You give it a photo of a person and a garment image; it dresses the person in that garment while preserving pose, body shape, and the garment's pattern and text. Good for showing a clothing product on a model for an ecommerce listing.

    $0.028 / run
    productvirtual-try-onvton
  • Image 01

    ImageMiniMax
    New

    Minimax's first image model, with character reference support

    $0.012 / run
    replicateminimaxtext-to-image
  • Image 3.2

    ImageBria AI
    New

    Commercial-ready, trained entirely on licensed data, text-to-image model. With only 4B parameters provides exceptional aesthetics and text rendering. Evaluated to be on par to other leading models in the market

    $0.048 / run
    replicatebriatext-to-image
  • InstructPix2Pix

    ImageCommunity

    Berkeley InstructPix2Pix. Edits an image from natural-language instructions in a single forward pass. Trained on GPT-3 plus Stable Diffusion synthetic pairs.

    $0.0039 / run
    replicatestyle-transferimage-edit
  • Tencent's face-identity conditioning adapter for SD/SDXL. Face embedding + CLIP for ID-consistent generation.

    $0.028 / run
    tencentimage-editface-id
  • Janus Pro 7B

    ImageDeepSeek

    DeepSeek's unified multimodal model. Decouples vision encoding for both understanding and generation tasks.

    $0.017 / run
    deepseekjanusopen-weights
  • Kling V2.5 Turbo Pro

    VideoKuaishou (Kling)
    New

    Kling 2.5 Turbo Pro: Unlock pro-level text-to-video and image-to-video creation with smooth motion, cinematic depth, and remarkable prompt adherence.

    $0.42 / run
    replicatekwaivgitext-to-video
  • Kokoro TTS 82M

    Text-to-SpeechCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    $0.00030 / run
    kokorottsopen-weights
  • Kuaishou Kolors

    ImageCommunity

    Kuaishou's bilingual (CN/EN) latent diffusion text-to-image model with strong text rendering.

    $0.089 / run
    kuaishoutext-to-imageopen-weights
  • LivePortrait

    VideoCommunity

    Kuaishou LivePortrait. Efficient portrait animation driven by reference videos with stitching, retargeting and motion-control parameters.

    $0.095 / run
    replicatelipsynckuaishou
  • Meta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.

    $0.0039 / run
    replicatemetallama
  • Llama 3.2 Vision 90B

    MultimodalCommunity

    Meta Llama 3.2 90B Vision. Largest open-weights Llama vision model. Strong visual reasoning, chart, OCR and document understanding.

    $0.0070 / run
    replicatemultimodalvision-understanding
  • LLaVA 1.6 Vicuna 13B

    MultimodalCommunity

    LLaVA 1.6 (LLaVA-NeXT) with a Vicuna-13B language backbone. Open vision-language chat model that describes images, answers questions, reads charts and reasons about scenes. Version 1.6 adds higher input resolution and better OCR and reasoning than LLaVA 1.5.

    $0.10 / run
    replicatellavacaptioning
  • SDXL fine-tune by mejiabrayan aimed at logo generation. Produces simple, centered mark and wordmark style logos from a text prompt. Useful for quick brand concepts and mockups. Raster PNG output, not vector.

    $0.037 / run
    replicatelogoicon
  • Lotus-G

    MultimodalCommunity

    Lotus generative depth model. Treats depth as a generation task using a diffusion model, producing higher-fidelity depth on textured surfaces.

    $0.041 / run
    replicatedepthvision-understanding
  • LTX-Video (Lightricks)

    VideoLightricks

    Lightricks' 2B DiT video model. Realtime generation on consumer GPUs (~6s @ H100, 24fps).

    $0.023 / run
    lightricksltxtext-to-video
  • Lucid Origin

    ImageCommunity
    New

    Artistic and high-quality visuals with improved prompt adherence, diversity, and definition

    $0.020 / run
    replicateleonardoaitext-to-image
  • Luma Ray Flash 2

    VideoLuma AI
    New

    Fast affordable video with I2V support

    $0.36 / run
    fastbudgeti2v
  • Luma Ray-2 720p

    VideoLuma AI

    Luma Labs' Ray-2 at 720p on Replicate. Text and image-to-video producing 5s and 9s clips with fast, coherent motion and strong camera control. Successor to Dream Machine.

    $1.08 / run
    replicatelumaray-2
  • MagicAnimate

    VideoCommunity

    ByteDance MagicAnimate. Temporally consistent human-image animation driven by a DensePose motion sequence with strong identity preservation.

    $0.41 / run
    replicateanimationhuman-motion
  • Magicoder S CL 7B

    CodeCommunity

    UIUC Magicoder S CL 7B. CodeLlama-7B fine-tuned with OSS-Instruct synthetic data. Strong HumanEval Plus and MBPP Plus performance per parameter.

    $0.092 / run
    replicatecode-generationopen-weights
  • MAGNeT

    Audio & MusicCommunity

    MAGNeT is Meta's masked, non-autoregressive audio generator. Instead of predicting tokens left to right it fills masked audio tokens in parallel over a few decoding steps, so generation is faster than autoregressive MusicGen at similar quality. This Replicate packaging exposes the text-to-music and text-to-sound variants.

    $0.0030 / run
    metamagnetnon-autoregressive
  • Detail-hallucinating upscaler in the Magnific style. Adds plausible high-frequency texture using a Stable Diffusion refiner conditioned on the low-res input.

    $0.059 / run
    replicateupscalingcreative
  • Marigold

    MultimodalCommunity

    ETH Zurich Marigold. Diffusion-based monocular depth-estimation model fine-tuned from Stable Diffusion with strong fine-detail recovery.

    $0.094 / run
    replicatedepthvision-understanding
  • Mask2Former

    MultimodalMeta

    Meta Mask2Former universal image-segmentation transformer. Single architecture for panoptic, instance and semantic segmentation tasks.

    $0.028 / run
    replicatesegmentationvision-understanding
  • MiDaS v3.1

    MultimodalCommunity

    Intel MiDaS v3.1 relative depth-estimation model. Robust zero-shot single-image depth across diverse domains and resolutions.

    $0.00030 / run
    replicatedepthvision-understanding
  • MiniCPM-V 2.6

    MultimodalCommunity

    OpenBMB MiniCPM-V 2.6. 8B vision-language model with strong single-image, multi-image and video understanding plus OCR capabilities.

    $0.0015 / run
    replicatemultimodalvision-understanding
  • Minimax Video

    VideoMiniMax

    MiniMax's video generation model. Fast, high-quality video output with text-to-video capabilities.

    $0.60 / run
    fastaffordable
  • Mochi 1

    VideoCommunity

    Genmo's 10B open-weights text-to-video model. AsymmDiT architecture, 5.4s @ 480p.

    $0.50 / run
    genmomochitext-to-video
  • Mochi 1

    VideoGenmo

    Genmo's Mochi 1, an open text-to-video model with high-fidelity motion built on a 10B Asymmetric Diffusion Transformer. Released under Apache 2.0, it was the largest open video model at launch and is strong on smooth, physically plausible movement.

    $0.50 / run
    replicategenmomochi
  • Molmo 7B

    MultimodalCommunity

    Allen AI Molmo 7B-D on Replicate. Open vision-language model trained on the PixMo data, notable for pointing at and locating objects in images, not just describing them.

    $0.050 / run
    replicateallenaimolmo
  • Moondream2

    MultimodalCommunity

    Moondream2 small vision-language model on Replicate. About 1.9B params, designed to run on edge devices, handles captioning, visual QA and short OCR-style reads at very low cost.

    $0.0020 / run
    replicatemoondreamvision-understanding
  • MuseTalk

    VideoCommunity

    Tencent MuseTalk real-time lip-sync model. Audio-driven mouth-region editing in latent space at 30+ fps on a single GPU.

    $0.062 / run
    replicatelipsynctencent
  • Nano Banana

    ImageGoogle DeepMind
    New

    Google's latest image editing model in Gemini 2.5

    $0.047 / run
    replicategoogletext-to-image
  • Nano Banana 2

    ImageGoogle DeepMind
    New

    Google's fast image generation model with conversational editing, multi-image fusion, and character consistency

    $0.080 / run
    replicategoogletext-to-image
  • Nano Banana 2 Lite

    ImageGoogle DeepMind
    New

    Google's fastest image generation model — the lightweight, low-cost version of Nano Banana 2, for rapid creation and editing

    $0.041 / run
    replicategoogletext-to-image
  • Nano Banana Pro

    ImageGoogle DeepMind
    New

    Google's state-of-the-art image generation and editing model (Gemini 3 Pro Image): sharp text rendering, 1K to 4K output.

    $0.18 / run
    replicategoogletext-to-image
  • olmOCR

    MultimodalCommunity

    Allen AI olmOCR. Open-source 7B vision-language model fine-tuned for high-fidelity document parsing including math, code and tables.

    $0.020 / run
    replicateocrvision-understanding
  • OOTDiffusion (Try-On)

    ImageCommunity

    OOTDiffusion virtual try-on. Takes a clear photo of a model and an upper-body garment and renders the garment onto the person using an outfitting-fusion diffusion approach that keeps the garment's texture and the model's pose. A lightweight alternative to IDM-VTON for clothing previews.

    $0.18 / run
    productvirtual-try-onvton
  • OpenPose

    MultimodalCommunity

    CMU OpenPose multi-person 2D pose estimator. Real-time keypoint detection for body, hand, face and foot using Part Affinity Fields.

    $0.0012 / run
    replicateposevision-understanding
  • OpenVoice v2

    Text-to-SpeechCommunity

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    $0.067 / run
    myshellttsvoice-cloning
  • P Image

    ImagePruna AI
    New

    A sub 1 second text-to-image model built for production use cases.

    $0.0060 / run
    replicateprunaaitext-to-image
  • PaddleOCR v3

    MultimodalCommunity

    Baidu PaddleOCR v3 PP-OCR pipeline. Lightweight detector plus recognizer optimized for production use with 80+ language support.

    $0.022 / run
    replicateocrvision-understanding
  • Parler-TTS

    Text-to-SpeechCommunity

    Hugging Face Parler-TTS Mini. Lightweight TTS conditioned on a natural-language style description for fine-grained control over voice characteristics.

    $0.0024 / run
    parlerttshuggingface
  • Phind CodeLlama 34B v2. Highly tuned CodeLlama variant focused on retrieval-augmented developer assistant workflows.

    $0.22 / run
    replicatecode-generationphind
  • PhotoMaker

    ImageTencent ARC

    Tencent ARC PhotoMaker. Identity-preserving stylized photo generation from a stacked-ID embedding. Realistic re-styling of a subject in seconds.

    $0.0080 / run
    replicatestyle-transferimage-edit
  • Photon

    ImageLuma AI
    New

    High-quality image generation model optimized for creative professional workflows and ultra-high fidelity outputs

    $0.036 / run
    replicatelumatext-to-image
  • Photon Flash

    ImageLuma AI
    New

    Accelerated variant of Photon prioritizing speed while maintaining quality

    $0.012 / run
    replicatelumatext-to-image
  • PixVerse v5.6

    VideoPixVerse
    New

    Physics-accurate video generation up to 1080p

    $0.42 / run
    i2v1080pphysics
  • Playground AI's diffusion model tuned for aesthetics. SDXL-based architecture trained on the EDM formulation, rated by users as more visually pleasing than SDXL in their study. Strong on vivid color and contrast.

    $0.073 / run
    replicateplaygroundtext-to-image
  • Point-E

    ImageCommunity

    OpenAI Point-E text-to-point-cloud system. Fast 3D point-cloud generation from text, optionally lifted to a mesh via marching cubes.

    $0.071 / run
    replicate3d-generationopenai
  • Qwen Image

    ImageAlibaba (Qwen)
    New

    An image generation foundation model in the Qwen series that achieves significant advances in complex text rendering.

    $0.030 / run
    replicateqwentext-to-image
  • Qwen Image 2

    ImageAlibaba (Qwen)
    New

    A next-generation image generation and editing model from Alibaba's Qwen team. Supports text-to-image and image editing with strong text rendering, especially for Chinese.

    $0.042 / run
    replicateqwentext-to-image
  • Qwen Image 2 Pro

    ImageAlibaba (Qwen)
    New

    The pro version of Qwen Image 2 from Alibaba's Qwen team. Enhanced text rendering, realism, and semantic adherence for high-quality image generation and editing.

    $0.090 / run
    replicateqwentext-to-image
  • Qwen Image 2512

    ImageAlibaba (Qwen)
    New

    Qwen Image 2512 is an improved version of Qwen Image with more realistic human generation, finer textures, and stronger text rendering

    $0.024 / run
    replicateqwentext-to-image
  • Qwen Image 3

    ImageAlibaba (Qwen)
    New

    Qwen-Image-3.0 generates and edits images with accurate text rendering, complex layouts, and photographic detail.

    $0.036 / run
    replicatealibabatext-to-image
  • Qwen Image 3 Pro

    ImageAlibaba (Qwen)
    New

    Qwen-Image-3.0-Pro generates and edits images with dense, accurate text rendering, complex multi-element layouts, and photographic detail.

    $0.048 / run
    replicatealibabatext-to-image
  • Qwen-Image-Edit

    ImageAlibaba (Qwen)

    Alibaba Qwen's instruction-driven image editor. Extends Qwen-Image's text-rendering ability to editing, so it handles both semantic edits (swap objects, change style) and precise text edits inside the image while preserving the original layout and unedited regions.

    $0.036 / run
    qwenalibabaimage-edit
  • Qwen2-VL 7B Instruct

    MultimodalCommunity

    Alibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.

    $0.0024 / run
    replicateqwenalibaba
  • Qwen3 TTS

    Text-to-SpeechAlibaba (Qwen)
    New

    A unified Text-to-Speech demo featuring three powerful modes: Voice, Clone and Design

    $0.024 / run
    replicateqwentext-to-speech
  • Rd Animation

    ImageCommunity
    New

    Style consistent animated pixel art sprite generation

    $0.084 / run
    replicateretro-diffusiontext-to-image
  • Real-ESRGAN 4x

    ImageCommunity

    AI-Upscaler that increases image resolution up to 4x while preserving texture and detail. Trained on synthetic and real data to reduce common ESRGAN artifacts.

    $0.0024 / run
    replicateupscalingimage-restore
  • Real-ESRGAN Anime 4x

    ImageCommunity

    Real-ESRGAN variant fine-tuned for anime, manga, and illustrated artwork. 4x upscaling with cartoon-aware artifact suppression.

    $0.0051 / run
    replicateupscalinganime
  • Recraft 20B

    ImageRecraft
    New

    Affordable and fast images

    $0.026 / run
    replicaterecraft-aitext-to-image
  • Recraft V3

    ImageRecraft
    New

    State-of-the-art image generation optimized for design and branding. SVG vector output support.

    $0.048 / run
    designvectorbranding
  • Recraft v3 SVG

    ImageRecraft

    Recraft's v3 variant that outputs vector SVG instead of raster pixels. Generates clean, editable logos, icons and illustrations that scale without quality loss, which is unusual among image models. Hosted API only.

    $0.096 / run
    replicaterecrafttext-to-image
  • Recraft V4

    ImageRecraft
    New

    Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and cost-efficient at standard resolution.

    $0.048 / run
    replicaterecraft-aitext-to-image
  • Recraft V4 Pro

    ImageRecraft
    New

    Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4, with higher resolution for print-ready and large-scale work.

    $0.30 / run
    replicaterecraft-aitext-to-image
  • Recraft V4 SVG

    ImageRecraft

    Recraft V4 SVG turns a text prompt into production-ready SVG vector art with clean geometry and structured, editable layers. Newer generation than V3 with improved design quality on logos, icons and flat illustration. Returns true vector paths, not a traced bitmap.

    $0.096 / run
    svgvectorrecraft
  • Recraft V4.1

    ImageRecraft
    New

    Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and cost-efficient at standard resolution.

    $0.048 / run
    replicaterecraft-aitext-to-image
  • Recraft V4.1 Pro

    ImageRecraft
    New

    Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4.1, with higher resolution for print-ready and large-scale work.

    $0.30 / run
    replicaterecraft-aitext-to-image
  • New

    A faster, lighter Recraft image generation model optimized for high-volume and production pipelines. Same design taste as V4.1, built for speed and throughput.

    $0.048 / run
    replicaterecraft-aitext-to-image
  • New

    A faster, lighter Recraft image generation model at ~2048px resolution, optimized for high-volume production. Design taste and prompt accuracy at high resolution with better throughput.

    $0.30 / run
    replicaterecraft-aitext-to-image
  • Rembg

    ImageCommunity

    Open-source background-removal tool wrapping U2Net. Produces alpha mattes for photos, products and people with no manual masking.

    $0.0047 / run
    replicatebackground-removalmatting
  • Lucataco's remove-bg, a rembg-based background removal model that returns the foreground subject on a transparent background. A popular, low-cost option for quick product and portrait cutouts.

    $0.00060 / run
    background-removallucatacorembg
  • Remove Object (LaMa)

    ImageCommunity

    Object removal and cleanup using LaMa inpainting. Paint a mask over an unwanted object, logo or person and the model fills the area with plausible background, erasing it from the photo.

    $0.00090 / run
    background-removalobject-removallama
  • Real-Time Intermediate Flow Estimation. Doubles or quadruples FPS of an existing video via learned optical-flow-based frame interpolation.

    $0.044 / run
    replicateupscaleframe-interpolation
  • Riffusion

    Text-to-SpeechRiffusion

    Stable-Diffusion-based real-time music generator. Operates on spectrogram images then resynthesizes audio, enables seamless transitions and looping.

    $0.058 / run
    riffusionmusic-generationopen-weights
  • RVC Voice Conversion

    Text-to-SpeechCommunity

    Retrieval-based Voice Conversion. Converts a source recording into a target speaker's voice, preserving pitch, prosody and rhythm.

    $0.050 / run
    rvcvoice-conversionvoice-cloning
  • SadTalker

    VideoCommunity

    Stylized audio-driven talking-head generator. Synthesizes 3D motion coefficients from audio to animate a single portrait image with natural head movements.

    $0.12 / run
    replicatelipsynctalking-head
  • SDXL Emoji

    ImageCommunity

    SDXL fine-tune by fofr trained on Apple emoji art. Generates rounded, glossy emoji and icon style graphics from a text prompt, useful for custom reactions, app glyphs and playful icon sets. Raster output.

    $0.016 / run
    replicateemojiicon
  • SDXL Inpainting

    ImageCommunity

    SDXL inpainting built on the Hugging Face Diffusers inpaint pipeline. Replace or remove masked regions of an image with prompt-conditioned content at SDXL resolution. A cheap, well-understood baseline for object removal and local edits.

    $0.0026 / run
    sdxlstability-aiimage-edit
  • SeamlessM4T

    Speech-to-TextCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    $0.16 / run
    replicatemetaseamless
  • SeamlessM4T v2 Large (Text)

    Text & ChatCommunity

    Meta SeamlessM4T v2 Large. Universal multilingual translation across 100+ languages with text-to-text mode for documents and chat.

    $0.0012 / run
    replicatetranslationmeta
  • Seedance 1 Pro Fast

    VideoByteDance
    New

    A faster and cheaper version of Seedance 1 Pro

    $0.36 / run
    replicatebytedancetext-to-video
  • Seedance Lite

    VideoByteDance
    New

    Budget ByteDance video, fast and cheap

    $0.22 / run
    budgeti2vfast
  • Seedance Pro

    VideoByteDance
    New

    ByteDance video with T2V and I2V, up to 1080p

    $0.90 / run
    i2v1080p
  • Seedream 4

    ImageByteDance
    New

    ByteDance's unified text-to-image generation and editing model with up to 4K resolution.

    $0.036 / run
    replicatebytedancetext-to-image
  • Seedream 4.5

    ImageByteDance
    New

    ByteDance Seedream 4.5: upgraded image generation and editing with stronger spatial understanding and world knowledge.

    $0.048 / run
    replicatebytedancetext-to-image
  • Seedream 5 Lite

    ImageByteDance
    New

    ByteDance Seedream 5 Lite: fast, affordable image generation and editing at 2K or 3K resolution.

    $0.042 / run
    replicatebytedancetext-to-image
  • Segformer B5

    MultimodalCommunity

    NVIDIA SegFormer-B5 semantic segmentation. Hierarchical transformer encoder with lightweight MLP decoder, strong ADE20k and Cityscapes results.

    $0.00030 / run
    replicatesegmentationvision-understanding
  • Shap-E (OpenAI)

    ImageCommunity

    OpenAI Shap-E text/image to 3D. Generates implicit neural representations renderable as textured meshes or NeRFs.

    $0.11 / run
    replicate3d-generationopenai
  • Spark TTS

    Text-to-SpeechCommunity

    Spark efficient TTS with disentangled control over speaker, content and style. Strong cross-lingual zero-shot performance.

    $0.0035 / run
    sparkttsvoice-cloning
  • Stable Audio Open 1.0

    Audio & MusicCommunity

    Stability AI's Stable Audio Open generates short audio from text prompts, tuned for sound effects, drum loops, instrument riffs and production elements rather than full songs. Open weights, latent diffusion over a 44.1kHz audio autoencoder, with a configurable seconds_total up to about 47 seconds.

    $0.0069 / run
    stability-aistable-audiosound-effects
  • Stable Diffusion 3

    ImageStability AI
    New

    A text-to-image model with greatly improved performance in image quality, typography, complex prompt understanding, and resource-efficiency

    $0.042 / run
    replicatestability-aitext-to-image
  • Stability AI's 8B MMDiT-based flagship. Open weights at 1MP with improved typography and prompt adherence over SDXL. The largest model in the SD 3.5 release line.

    $0.078 / run
    replicatestability-aistable-diffusion
  • Distilled, 4-step version of SD 3.5 Large from Stability AI. Keeps most of the large model's quality and text rendering at a fraction of the inference time. Open weights under the Stability Community License.

    $0.048 / run
    replicatestability-aistable-diffusion
  • New

    2.5 billion parameter image model with improved MMDiT-X architecture

    $0.042 / run
    replicatestability-aitext-to-image
  • Stable Diffusion XL

    ImageStability AI

    Stability AI's SDXL model via Replicate. High-quality image generation with extensive customization.

    $0.0015 / run
    open-sourcecustomizable
  • StarCoder2 15B

    CodeCommunity

    BigCode StarCoder2 15B code-generation flagship. Trained on 4T tokens of Stack v2 data with grouped-query attention and 16k context.

    $1.19 / run
    replicatecode-generationbigcode
  • StarVector 8B is a multimodal model that generates SVG code directly from an input image. Rather than tracing pixels, it predicts the SVG markup token by token, which can produce compact, semantically structured paths for icons and simple graphics. Research model from the StarVector project.

    $0.070 / run
    svgvectorstarvector
  • StyleTTS 2

    Text-to-SpeechCommunity

    Style-based TTS using diffusion and adversarial training. Human-level naturalness in zero-shot voice synthesis from a 3-5s reference clip.

    $0.00040 / run
    stylettsttsvoice-cloning
  • Suno Bark

    Text-to-SpeechSuno

    Suno's text-prompted generative audio model. Speech, music, ambient sound and effects with non-verbal cues like laughter or sighs.

    $0.097 / run
    sunobarkmusic-generation
  • SUPIR

    ImageCommunity

    SUPIR by Fanghua Yu et al. is a large diffusion-based restoration model that recovers photorealistic detail from heavily degraded images and can be steered with a text prompt describing the scene.

    $0.49 / run
    restoresuper-resolutionsupir
  • SUPIR Upscaler

    ImageCommunity

    SUPIR (Scaling-Up Image Restoration) photo-real restoration model. Combines SDXL prior with language-guided controls for severely degraded inputs.

    $0.30 / run
    replicateupscalingimage-restore
  • Swin2SR

    ImageCommunity

    Transformer-based image super-resolution using Swin-V2 attention. Handles classical, lightweight, real-world, and compressed-input variants with 2x/4x upscaling.

    $0.0088 / run
    replicateupscalingtransformer
  • SwinIR Video

    VideoCommunity

    SwinIR transformer-based super-resolution and denoising applied per-frame to video. Handles classic, real-world and lightweight upscaling.

    $0.028 / run
    replicateupscaletransformer
  • ToonCrafter

    VideoCommunity

    Tencent ToonCrafter generative cartoon interpolation model. Synthesizes smooth in-between frames between two cartoon keyframes.

    $0.085 / run
    replicateanimationtooncrafter
  • Tortoise TTS

    Text-to-SpeechCommunity

    Multi-voice expressive TTS. Slow but high-quality with strong prosody and natural intonation. Trained for long-form narration use cases.

    $0.082 / run
    tortoisettsexpressive
  • TRELLIS (3D)

    ImageCommunity

    Microsoft TRELLIS image-to-3D model. Generates textured 3D assets in GLB or Gaussian-splat format from a single reference image.

    $0.042 / run
    replicate3d-generationimage-to-3d
  • V-Express

    VideoCommunity

    Tencent V-Express. Audio-driven portrait animation with progressive training, weak-condition learning, and expressive lip sync.

    $0.10 / run
    replicatelipsynctencent
  • Vectorizer (VTracer)

    ImageCommunity

    PNG/JPG to SVG vectorizer built on VTracer, the open-source raster-to-vector engine. Traces a bitmap into layered color regions and clean paths with controls for color count, area threshold and path simplification. Fast, deterministic alternative to model-based vectorizers.

    $0.0022 / run
    svgvectorvtracer
  • VideoCrafter

    VideoCommunity

    Tencent VideoCrafter latent video diffusion. Text-to-video and image-to-video generation up to 2s at 1024x576 with strong motion fidelity.

    $0.16 / run
    replicateupscalevideo-generation
  • Wan 2.2 5B Fast

    VideoAlibaba (Wan)
    New

    The fastest Wan 2.2 text-to-image and image-to-video model

    $0.030 / run
    replicatewan-videotext-to-video
  • Wan 2.2 Image

    ImageCommunity
    New

    This model generates beautiful cinematic 2 megapixel images in 3-4 seconds and is derived from the Wan 2.2 model through optimisation techniques from the pruna package

    $0.024 / run
    replicateprunaaitext-to-image
  • Wan 2.2 Image-to-Video

    VideoAlibaba (Wan)
    New

    Ultra-cheap I2V. Upload image and animate it.

    $0.060 / run
    budgeti2vfast
  • Wan 2.7 Image

    ImageAlibaba (Wan)
    New

    Generate and edit images with Alibaba's Wan 2.7

    $0.036 / run
    replicatewan-videotext-to-image
  • Wan 2.7 Image Pro

    ImageAlibaba (Wan)
    New

    Generate and edit high-quality images with Alibaba's Wan 2.7 Pro with 4K output, thinking mode, text-to-image, multi-image editing, and image set generation

    $0.036 / run
    replicatewan-videotext-to-image
  • Wan 3

    VideoAlibaba (Qwen)
    New

    (50% off until Aug 30!) - Wan 3.0 generates video from a text prompt or a starting image, with cinematic motion and support for 480p, 720p, and 1080p output up to 30 seconds.

    $0.60 / run
    replicatealibabatext-to-video
  • Wav2Lip

    VideoCommunity

    Lip-sync model that re-syncs a target video's lip movement to an arbitrary audio track. Robust to identity and language with a lip-sync discriminator loss.

    $0.0067 / run
    replicatelipsyncvideo-edit
  • Whisper Diarization

    Speech-to-TextCommunity

    Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

    $0.0063 / run
    replicatewhisperstt
  • WhisperX

    Speech-to-TextCommunity

    WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

    $0.024 / run
    replicatewhisperxstt
  • WizardCoder 33B

    CodeCommunity

    WizardLM WizardCoder 33B v1.1. Evol-Instruct fine-tune of DeepSeek-Coder-33B with strong code-generation benchmark performance.

    $0.025 / run
    replicatecode-generationwizardlm
  • ZoeDepth

    MultimodalCommunity

    Intel ZoeDepth metric depth-estimation model. Combines relative-depth pretraining with metric fine-tuning for absolute distance in real units.

    $0.00060 / run
    replicatedepthvision-understanding

Older versions (1)

Still runnable, superseded by a newer model. Not counted above, and hidden by default in the catalog.

Currently unavailable (9)

Listed for reference. These models have no verified price or are no longer served, so the API does not run them.

Frequently asked questions

How is Replicate pricing handled on Railwail?
Railwail uses transparent per-call or per-token credit pricing for all Replicate models. You pay only for what you use — no monthly minimums, no upfront commitments. Pricing for every individual Replicate model is shown on its detail page.
Are there rate limits when using Replicate via Railwail?
Default rate limits depend on your account tier and the underlying Replicate capacity. Free-tier accounts get sensible defaults for development; paid accounts can request higher limits. Contact support if you need dedicated throughput or burst capacity.
Which regions does Replicate support through Railwail?
Replicate models are served from Railwail's globally distributed edge infrastructure. EU, US, and Asia-Pacific traffic is automatically routed to the nearest available provider region. GDPR-compliant EU-only routing is available on request.
Is there a sandbox or free tier to test Replicate models?
Accounts created with Google get 10 trial credits ($0.10), usable 24 hours after sign-up for runs of up to 2 credits each, so you can test the cheaper Replicate models. Email and GitHub sign-ups start without trial credits. No credit card is required to sign up.
Categories Replicate works in

Start building with Replicate today

Free credits on sign-up. No credit card required. Access Replicate and 27+ other providers through a single API.