Replicate is an open-model hosting platform that serves thousands of open-source models including Flux, Stable Diffusion, Llama, and Whisper variants via a unified API.
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
Black Forest Labs' flagship text-to-image model. Faster generation than FLUX.1 Pro at higher prompt adherence, with strong photorealism and reliable spatial composition. Runs as a hosted Replicate model.
Google DeepMind's Imagen 4 text-to-image model, hosted on Replicate. Sharp detail, accurate text rendering, and strong prompt adherence across photographic and illustrated styles. Outputs up to 2K resolution.
Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.
SDXL fine-tune by galleri5 for slick flat icons and pop constructivist graphics with thick edges. Trained on Bing generations, it produces clean single-subject icon art that suits app icons, badges and UI glyphs. Raster output, not true vector.
The highest-quality tier of Ideogram v3. Improved photorealism and prompt adherence over v2 while keeping Ideogram's best-in-class text rendering. Supports style references and inline text layout.
Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.
InstantID makes realistic portraits of a real person from a single reference photo without per-user training. Combines a face encoder with an IdentityNet adapter on SDXL to keep identity and pose while following a text prompt, so it is fast and tuning-free.
Kuaishou's Kling v2.1, generating 5 and 10 second videos at 720p or 1080p from text or an image. Known for cinematic camera work and realistic physical motion, available on Replicate via the official KwaiVGI account.
Kuaishou's premium Kling v2.1 Master. Generates 1080p 5s and 10s clips from text or an image with strong dynamics and prompt adherence. The top tier of the Kling 2.1 family.
MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.
Turns any single selfie into a clean professional headshot using FLUX Kontext image editing. Keeps the person's face while swapping to business attire, a studio background and even lighting. Aimed at LinkedIn-style profile photos.
Recraft's faster, cheaper vector model. Outputs editable SVG paths instead of raster pixels, so logos, icons and flat illustrations scale to any size without blur. Defaults to a vector_illustration style and supports line art and engraving looks. Hosted API only.
Recraft's raster-to-vector converter. Takes a PNG or JPG and traces it into a clean SVG with precise vector paths, aimed at logos, icons and graphics that need to scale. Image-to-SVG counterpart to Recraft's text-to-SVG models.
Runway's Gen-4 Turbo on Replicate. Fast image-to-video generation producing 5s and 10s clips at 720p with strong character and scene consistency across shots.
Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.
fofr's sticker generator that outputs graphics with transparent backgrounds, so the result drops straight into chat apps or print sheets. Runs an SDXL-based pipeline at high speed (default 17 steps) and returns die-cut style art without manual background removal.
OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.
Background removal model from 851-Labs that outputs a clean cutout with a transparent alpha channel. One of the most-run background removers on Replicate, handles people, products and objects on busy backgrounds.
Product advertising photo generator. You upload a cut-out product shot and a prompt describing the scene; it places the product on a new generated background with matching lighting and shadows, so a plain packshot becomes an ecommerce or ad-ready hero image without a photo studio.
BiRefNet high-resolution dichotomous image segmentation for background removal. Bilateral reference network that produces sharp matting on fine detail like hair, fur and thin structures, often cleaner than older U2Net or rembg models.
BRIA AI's commercial background removal model trained on fully licensed data. Produces accurate cutouts for e-commerce and design, with attention to clean edges around products and people.
Microsoft Research pipeline by Ziyu Wan et al. that restores scanned old photos, removing scratches, dust and fading and optionally enhancing faces in one pass.
catacolabs Cartoonify turns a photo into a flat cartoon illustration. Takes a single image and returns a stylized cartoon version with clean shapes and bold outlines. Straightforward one-input model for avatars and profile pictures.
Resemble AI's open Chatterbox TTS. Zero-shot voice cloning from a short audio prompt with an exaggeration control for emotion intensity, plus CFG weight to balance pacing and fidelity.
Meta's 13B Code Llama tuned for instruction following. A faster mid-size option for code generation and completion, supporting infilling for inserting code at a cursor position. Served on Replicate per call.
Meta's 34B Code Llama tuned for instruction following. A balance of size and quality for code generation, completion, and explanation, with strong coverage of Python, JavaScript, and other common languages. Runs on Replicate per call.
Meta's largest Code Llama, a 70B Llama-2 derivative specialized for programming and tuned to follow instructions in chat form. Handles code generation, completion, and explanation across common languages. Served on Replicate as a per-call endpoint.
Meta's smallest Code Llama at 7B parameters, tuned for instruction following. The cheapest and fastest member of the family for quick code generation, completion, and infilling. Served on Replicate per call.
Robust face-restoration model using a transformer-based codebook prior. Handles severe degradation, occlusion, and old-photo restoration with adjustable fidelity-quality tradeoff.
CogVideoX-5B from Tsinghua/Zhipu AI, an open 5B-parameter text-to-video diffusion transformer. Generates 6-second 720p clips with coherent motion and is widely used in research for its open weights and reproducibility.
fofr's model generates the same character in many poses and angles from one reference image. Useful for building an avatar set or character sheet where the face and design stay consistent across outputs. Can produce a grid or individual images.
DDColor by Xiaoyang Kang et al. colorizes black-and-white photos using dual decoders that jointly learn pixel colors and semantic color queries, giving vivid and natural results on old images.
Quantized GGUF build of DeepSeek's 33B code model, trained on roughly 2T tokens that are about 87 percent code. Designed for repository-level completion and project-aware generation thanks to a 16k context window. Runs on Replicate as a per-call endpoint.
fofr's face stylizer converts a face photo into 3D render, emoji, pixel art, video-game character, claymation or toy styles. Uses InstantID plus style LoRAs on SDXL to keep the likeness while applying a chosen art style. Popular for fun avatars.
fofr's model turns a face photo into a die-cut sticker with a white border and transparent background. Uses InstantID to hold the likeness and outputs a clean PNG suitable for chat stickers or print. Simple single-image input.
SOTA Open source model trained on licensed data, transforming intent into structured control for precise, high-quality AI image generation in enterprise and agentic workflows.
Microsoft Florence-2 Large. Unified prompt-based vision foundation model for captioning, detection, segmentation and OCR with a single 770M-param backbone.
A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural language prompts
A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming images through natural language
An opinionated text-to-image model from Black Forest Labs in collaboration with Krea that excels in photorealism. Creates images that avoid the oversaturated "AI look".
PuLID identity customization running on FLUX.1-dev. Inserts a face from one reference photo into prompt-driven scenes using contrastive alignment, giving higher likeness and detail than SDXL-era ID adapters. Good for realistic avatars and character portraits.
The open-weight 12B rectified-flow transformer from Black Forest Labs. Close to FLUX Pro quality with a guidance-distilled checkpoint released under a non-commercial license. The most widely fine-tuned base in the FLUX family.
The fastest FLUX model from Black Forest Labs, distilled to produce images in 1 to 4 steps. Apache 2.0 licensed for commercial use. Built for high-volume generation and real-time previews.
Open-weight edge-guided FLUX model from Black Forest Labs. Extracts Canny edges from a control image and regenerates it from your prompt while holding the original composition and outlines, so you can restyle a scene without changing its structure.
Open-weight depth-guided FLUX model from Black Forest Labs. Derives a depth map from the control image and regenerates from your prompt while preserving 3D spatial layout, useful for re-texturing rooms, products, or scenes without moving objects.
Black Forest Labs' open-weight inpainting and outpainting model, guidance-distilled from FLUX.1 Fill [pro]. You supply an image plus a mask and a prompt; it fills the masked region or extends the canvas with prompt-guided content that matches lighting and texture.
Open-weight version of FLUX.1 Kontext by Black Forest Labs. Instruction-based editing: pass an input image and a plain text edit ('change the jacket to red', 'remove the person on the left') and it applies the change while keeping the rest of the scene and identity consistent.
FLUX.1-dev inpainting wrapper that fills masked parts of an image from a prompt. Useful when you want FLUX-quality fills with a simple image plus mask plus prompt interface and adjustable mask strength.
OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.
OpenAI's fastest model for high-quality, everyday image generation. Generate and edit images from text and image inputs with strong instruction following and sharp text rendering.
OpenAI's most capable image model, built for workflows where editing precision matters most. Generate and edit images from text and image inputs with strong instruction following, sharp text rendering, and detailed control.
Lvmin Zhang's IC-Light packaged by zsxkib. Relights a product or portrait from a text prompt or a chosen light direction while keeping the subject's shape and detail intact, so a flat product photo can be given studio, window, or dramatic side lighting without re-shooting.
Ideogram's text-to-image model known for accurate in-image text and typography. Handles posters, logos, and signage where other models garble lettering. Supports magic prompt expansion and multiple aspect ratios.
Ideogram's fast v3 model, the fastest and cheapest tier of the v3 family. Known for accurate in-image text rendering and reliable typography, which most diffusion models still get wrong. Hosted API only.
IDM-VTON virtual try-on from the CVPR 2024 paper. You give it a photo of a person and a garment image; it dresses the person in that garment while preserving pose, body shape, and the garment's pattern and text. Good for showing a clothing product on a model for an ecommerce listing.
Commercial-ready, trained entirely on licensed data, text-to-image model. With only 4B parameters provides exceptional aesthetics and text rendering. Evaluated to be on par to other leading models in the market
Berkeley InstructPix2Pix. Edits an image from natural-language instructions in a single forward pass. Trained on GPT-3 plus Stable Diffusion synthetic pairs.
Meta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.
LLaVA 1.6 (LLaVA-NeXT) with a Vicuna-13B language backbone. Open vision-language chat model that describes images, answers questions, reads charts and reasons about scenes. Version 1.6 adds higher input resolution and better OCR and reasoning than LLaVA 1.5.
SDXL fine-tune by mejiabrayan aimed at logo generation. Produces simple, centered mark and wordmark style logos from a text prompt. Useful for quick brand concepts and mockups. Raster PNG output, not vector.
Luma Labs' Ray-2 at 720p on Replicate. Text and image-to-video producing 5s and 9s clips with fast, coherent motion and strong camera control. Successor to Dream Machine.
MAGNeT is Meta's masked, non-autoregressive audio generator. Instead of predicting tokens left to right it fills masked audio tokens in parallel over a few decoding steps, so generation is faster than autoregressive MusicGen at similar quality. This Replicate packaging exposes the text-to-music and text-to-sound variants.
Detail-hallucinating upscaler in the Magnific style. Adds plausible high-frequency texture using a Stable Diffusion refiner conditioned on the low-res input.
Genmo's Mochi 1, an open text-to-video model with high-fidelity motion built on a 10B Asymmetric Diffusion Transformer. Released under Apache 2.0, it was the largest open video model at launch and is strong on smooth, physically plausible movement.
Allen AI Molmo 7B-D on Replicate. Open vision-language model trained on the PixMo data, notable for pointing at and locating objects in images, not just describing them.
Moondream2 small vision-language model on Replicate. About 1.9B params, designed to run on edge devices, handles captioning, visual QA and short OCR-style reads at very low cost.
OOTDiffusion virtual try-on. Takes a clear photo of a model and an upper-body garment and renders the garment onto the person using an outfitting-fusion diffusion approach that keeps the garment's texture and the model's pose. A lightweight alternative to IDM-VTON for clothing previews.
Playground AI's diffusion model tuned for aesthetics. SDXL-based architecture trained on the EDM formulation, rated by users as more visually pleasing than SDXL in their study. Strong on vivid color and contrast.
A next-generation image generation and editing model from Alibaba's Qwen team. Supports text-to-image and image editing with strong text rendering, especially for Chinese.
The pro version of Qwen Image 2 from Alibaba's Qwen team. Enhanced text rendering, realism, and semantic adherence for high-quality image generation and editing.
Alibaba Qwen's instruction-driven image editor. Extends Qwen-Image's text-rendering ability to editing, so it handles both semantic edits (swap objects, change style) and precise text edits inside the image while preserving the original layout and unedited regions.
Alibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.
AI-Upscaler that increases image resolution up to 4x while preserving texture and detail. Trained on synthetic and real data to reduce common ESRGAN artifacts.
Recraft's v3 variant that outputs vector SVG instead of raster pixels. Generates clean, editable logos, icons and illustrations that scale without quality loss, which is unusual among image models. Hosted API only.
Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and cost-efficient at standard resolution.
Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4, with higher resolution for print-ready and large-scale work.
Recraft V4 SVG turns a text prompt into production-ready SVG vector art with clean geometry and structured, editable layers. Newer generation than V3 with improved design quality on logos, icons and flat illustration. Returns true vector paths, not a traced bitmap.
Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and cost-efficient at standard resolution.
Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4.1, with higher resolution for print-ready and large-scale work.
A faster, lighter Recraft image generation model optimized for high-volume and production pipelines. Same design taste as V4.1, built for speed and throughput.
A faster, lighter Recraft image generation model at ~2048px resolution, optimized for high-volume production. Design taste and prompt accuracy at high resolution with better throughput.
Lucataco's remove-bg, a rembg-based background removal model that returns the foreground subject on a transparent background. A popular, low-cost option for quick product and portrait cutouts.
Object removal and cleanup using LaMa inpainting. Paint a mask over an unwanted object, logo or person and the model fills the area with plausible background, erasing it from the photo.
Stylized audio-driven talking-head generator. Synthesizes 3D motion coefficients from audio to animate a single portrait image with natural head movements.
SDXL fine-tune by fofr trained on Apple emoji art. Generates rounded, glossy emoji and icon style graphics from a text prompt, useful for custom reactions, app glyphs and playful icon sets. Raster output.
SDXL inpainting built on the Hugging Face Diffusers inpaint pipeline. Replace or remove masked regions of an image with prompt-conditioned content at SDXL resolution. A cheap, well-understood baseline for object removal and local edits.
Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.
Stability AI's Stable Audio Open generates short audio from text prompts, tuned for sound effects, drum loops, instrument riffs and production elements rather than full songs. Open weights, latent diffusion over a 44.1kHz audio autoencoder, with a configurable seconds_total up to about 47 seconds.
Stability AI's 8B MMDiT-based flagship. Open weights at 1MP with improved typography and prompt adherence over SDXL. The largest model in the SD 3.5 release line.
Distilled, 4-step version of SD 3.5 Large from Stability AI. Keeps most of the large model's quality and text rendering at a fraction of the inference time. Open weights under the Stability Community License.
StarVector 8B is a multimodal model that generates SVG code directly from an input image. Rather than tracing pixels, it predicts the SVG markup token by token, which can produce compact, semantically structured paths for icons and simple graphics. Research model from the StarVector project.
SUPIR by Fanghua Yu et al. is a large diffusion-based restoration model that recovers photorealistic detail from heavily degraded images and can be steered with a text prompt describing the scene.
PNG/JPG to SVG vectorizer built on VTracer, the open-source raster-to-vector engine. Traces a bitmap into layered color regions and clean paths with controls for color count, area threshold and path simplification. Fast, deterministic alternative to model-based vectorizers.
This model generates beautiful cinematic 2 megapixel images in 3-4 seconds and is derived from the Wan 2.2 model through optimisation techniques from the pruna package
Generate and edit high-quality images with Alibaba's Wan 2.7 Pro with 4K output, thinking mode, text-to-image, multi-image editing, and image set generation
(50% off until Aug 30!) - Wan 3.0 generates video from a text prompt or a starting image, with cinematic motion and support for 480p, 720p, and 1080p output up to 30 seconds.
Lip-sync model that re-syncs a target video's lip movement to an arbitrary audio track. Robust to identity and language with a lip-sync discriminator loss.
Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.
WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.
Railwail uses transparent per-call or per-token credit pricing for all Replicate models. You pay only for what you use — no monthly minimums, no upfront commitments. Pricing for every individual Replicate model is shown on its detail page.
Are there rate limits when using Replicate via Railwail?
Default rate limits depend on your account tier and the underlying Replicate capacity. Free-tier accounts get sensible defaults for development; paid accounts can request higher limits. Contact support if you need dedicated throughput or burst capacity.
Which regions does Replicate support through Railwail?
Replicate models are served from Railwail's globally distributed edge infrastructure. EU, US, and Asia-Pacific traffic is automatically routed to the nearest available provider region. GDPR-compliant EU-only routing is available on request.
Is there a sandbox or free tier to test Replicate models?
Accounts created with Google get 10 trial credits ($0.10), usable 24 hours after sign-up for runs of up to 2 credits each, so you can test the cheaper Replicate models. Email and GitHub sign-ups start without trial credits. No credit card is required to sign up.