Image Generation

Generate, edit, upscale or cut out images from a prompt or a source image.

Available
134of 162
Providers
20
Price range
$0.00050 – $0.4921per default run

133 models

  • FLUX 1.1 Pro

    Black Forest Labs

    Black Forest Labs' flagship text-to-image model. Faster generation than FLUX.1 Pro at higher prompt adherence, with strong photorealism and reliable spatial composition. Runs as a hosted Replicate model.

    fluxtext-to-imagephotorealism

    Modalities: Text, Image

    $0.048/image

  • Flux 1.1 Pro Ultra

    Black Forest Labs

    FLUX 1.1 Pro in ultra mode. Up to 4 megapixel images with raw mode for photorealism.

    high-qualityphotorealistic

    Modalities: Image

    $0.072/image

  • Flux Dev

    Black Forest Labs

    Black Forest Labs' development model. Fast, high-quality image generation with LoRA support.

    fastlora

    Modalities: Image

    $0.0384/image

  • Google Imagen 4

    Google DeepMind

    Google DeepMind's Imagen 4 text-to-image model, hosted on Replicate. Sharp detail, accurate text rendering, and strong prompt adherence across photographic and illustrated styles. Outputs up to 2K resolution.

    imagentext-to-imagephotorealism

    Modalities: Text, Image

    $0.048/image

  • SDXL fine-tune by galleri5 for slick flat icons and pop constructivist graphics with thick edges. Trained on Bing generations, it produces clean single-subject icon art that suits app icons, badges and UI glyphs. Raster output, not true vector.

    iconlogosdxl

    Modalities: Image

    β‰ˆ $0.0094/run

  • The highest-quality tier of Ideogram v3. Improved photorealism and prompt adherence over v2 while keeping Ideogram's best-in-class text rendering. Supports style references and inline text layout.

    text-to-imagetypographyphotorealism

    Modalities: Text, Image

    $0.108/image

  • InstantID

    Community

    InstantID makes realistic portraits of a real person from a single reference photo without per-user training. Combines a face encoder with an IdentityNet adapter on SDXL to keep identity and pose while following a text prompt, so it is fast and tuning-free.

    avatarportraitinstant-id

    Modalities: Text, Image

    β‰ˆ $0.084/run

  • Turns any single selfie into a clean professional headshot using FLUX Kontext image editing. Keeps the person's face while swapping to business attire, a studio background and even lighting. Aimed at LinkedIn-style profile photos.

    avatarportraitheadshot

    Modalities: Image

    $0.048/image

  • Recraft's faster, cheaper vector model. Outputs editable SVG paths instead of raster pixels, so logos, icons and flat illustrations scale to any size without blur. Defaults to a vector_illustration style and supports line art and engraving looks. Hosted API only.

    svgvectorlogo

    Modalities: Image

    $0.0528/image

  • Recraft's raster-to-vector converter. Takes a PNG or JPG and traces it into a clean SVG with precise vector paths, aimed at logos, icons and graphics that need to scale. Image-to-SVG counterpart to Recraft's text-to-SVG models.

    svgvectorvectorize

    Modalities: Image

    $0.012/image

  • Sticker Maker

    Community

    fofr's sticker generator that outputs graphics with transparent backgrounds, so the result drops straight into chat apps or print sheets. Runs an SDXL-based pipeline at high speed (default 17 steps) and returns die-cut style art without manual background removal.

    stickertransparentfofr

    Modalities: Image

    β‰ˆ $0.0055/run

  • Fibo

    Bria AI

    New

    SOTA Open source model trained on licensed data, transforming intent into structured control for precise, high-quality AI image generation in enterprise and agentic workflows.

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • Flux Fast

    Community

    New

    This is the fastest Flux endpoint in the world.

    prunaaitext-to-image

    Modalities: Text, Image

    $0.0060/image

  • Flux Kontext Max

    Black Forest Labs

    New

    A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural language prompts

    text-to-image

    Modalities: Text, Image

    $0.096/image

  • Flux Kontext Pro

    Black Forest Labs

    New

    A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming images through natural language

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • Flux Krea Dev

    Black Forest Labs

    New

    An opinionated text-to-image model from Black Forest Labs in collaboration with Krea that excels in photorealism. Creates images that avoid the oversaturated "AI look".

    text-to-image

    Modalities: Text, Image

    $0.030/image

  • Flux Pro

    Black Forest Labs

    New

    State-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity.

    text-to-image

    Modalities: Text, Image

    $0.066/image

  • New

    Runway's Gen-4 Image model with references. Use up to 3 reference images to create the exact image you need. Capture every angle.

    runwaymltext-to-image

    Modalities: Text, Image

    $0.096/image

  • New

    OpenAI's latest image generation model with better instruction following and adherence to prompts

    text-to-image

    Modalities: Text, Image

    $0.1632/image

  • New

    OpenAI's state-of-the-art image generation model. Create and edit images from text with strong instruction following, sharp text rendering, and detailed editing.

    text-to-image

    Modalities: Text, Image

    $0.1536/image

  • New

    OpenAI's fastest model for high-quality, everyday image generation. Generate and edit images from text and image inputs with strong instruction following and sharp text rendering.

    text-to-image

    Modalities: Text, Image

    $0.30/image

  • New

    OpenAI's most capable image model, built for workflows where editing precision matters most. Generate and edit images from text and image inputs with strong instruction following, sharp text rendering, and detailed control.

    text-to-image

    Modalities: Text, Image

    $0.30/image

  • New

    SOTA image model from xAI

    text-to-image

    Modalities: Text, Image

    $0.024/image

  • New

    xAI's Grok Imagine Image 2.0 β€” text-to-image generation and editing with a quality control and output up to 2k

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • New

    This is an optimised version of the hidream-l1 model using the pruna ai optimisation toolkit!

    prunaaitext-to-image

    Modalities: Text, Image

    $0.0060/image

  • Html To Image

    Community

    New

    Html To Image on Replicate (intelligent-utilities/html-to-image)

    intelligent-utilitiestext-to-image

    Modalities: Text, Image

    $0.0012/image

  • New

    Generate high-quality 2K resolution images from text prompts

    text-to-image

    Modalities: Text, Image

    $0.024/image

  • New

    A powerful native multimodal model for image generation (PrunaAI squeezed)

    text-to-image

    Modalities: Text, Image

    $0.096/image

  • Hunyuan3D 2.1

    Community

    New

    Refreshed Hunyuan3D 2.1 with improved texture fidelity and PBR-material support. Image-to-3D with textured GLB output.

    3d-generationimage-to-3dtencent

    Modalities: Image

    β‰ˆ $0.264/run

  • New

    A fast image model with state of the art inpainting, prompt comprehension and text rendering.

    ideogram-aitext-to-image

    Modalities: Text, Image

    $0.060/image

  • Ideogram V2A

    Ideogram

    New

    Like Ideogram v2, but faster and cheaper

    ideogram-aitext-to-image

    Modalities: Text, Image

    $0.048/image

  • New

    Like Ideogram v2 turbo, but now faster and cheaper

    ideogram-aitext-to-image

    Modalities: Text, Image

    $0.030/image

  • New

    Balance speed, quality and cost. Ideogram v3 creates images with stunning realism, creative designs, and consistent styles

    ideogram-aitext-to-image

    Modalities: Text, Image

    $0.072/image

  • New

    Balance speed, quality and cost. Ideogram v4 creates images with stunning realism, creative designs, and consistent styles

    ideogram-aitext-to-image

    Modalities: Text, Image

    $0.072/image

  • New

    The highest quality Ideogram v4 model. v4 creates images with stunning realism, creative designs, and consistent styles

    ideogram-aitext-to-image

    Modalities: Text, Image

    $0.12/image

  • Image 01

    MiniMax

    New

    Minimax's first image model, with character reference support

    text-to-image

    Modalities: Text, Image

    $0.012/image

  • Image 3.2

    Bria AI

    New

    Commercial-ready, trained entirely on licensed data, text-to-image model. With only 4B parameters provides exceptional aesthetics and text rendering. Evaluated to be on par to other leading models in the market

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • Lucid Origin

    Community

    New

    Artistic and high-quality visuals with improved prompt adherence, diversity, and definition

    leonardoaitext-to-image

    Modalities: Text, Image

    $0.0201/image

  • Nano Banana

    Google DeepMind

    New

    Google's latest image editing model in Gemini 2.5

    text-to-image

    Modalities: Text, Image

    $0.0468/image

  • Nano Banana 2

    Google DeepMind

    New

    Google's fast image generation model with conversational editing, multi-image fusion, and character consistency

    text-to-image

    Modalities: Text, Image

    $0.0804/image

  • Nano Banana 2 Lite

    Google DeepMind

    New

    Google's fastest image generation model β€” the lightweight, low-cost version of Nano Banana 2, for rapid creation and editing

    text-to-image

    Modalities: Text, Image

    $0.0408/image

  • Nano Banana Pro

    Google DeepMind

    New

    Google's state-of-the-art image generation and editing model (Gemini 3 Pro Image): sharp text rendering, 1K to 4K output.

    text-to-image

    Modalities: Text, Image

    $0.18/image

  • P Image

    Pruna AI

    New

    A sub 1 second text-to-image model built for production use cases.

    prunaaitext-to-image

    Modalities: Text, Image

    $0.0060/image

  • Photon

    Luma AI

    New

    High-quality image generation model optimized for creative professional workflows and ultra-high fidelity outputs

    text-to-image

    Modalities: Text, Image

    $0.036/image

  • New

    Accelerated variant of Photon prioritizing speed while maintaining quality

    text-to-image

    Modalities: Text, Image

    $0.012/image

  • Qwen Image

    Alibaba (Qwen)

    New

    An image generation foundation model in the Qwen series that achieves significant advances in complex text rendering.

    text-to-image

    Modalities: Text, Image

    $0.030/image

  • Qwen Image 2

    Alibaba (Qwen)

    New

    A next-generation image generation and editing model from Alibaba's Qwen team. Supports text-to-image and image editing with strong text rendering, especially for Chinese.

    text-to-image

    Modalities: Text, Image

    $0.042/image

  • Qwen Image 2 Pro

    Alibaba (Qwen)

    New

    The pro version of Qwen Image 2 from Alibaba's Qwen team. Enhanced text rendering, realism, and semantic adherence for high-quality image generation and editing.

    text-to-image

    Modalities: Text, Image

    $0.090/image

  • Qwen Image 2512

    Alibaba (Qwen)

    New

    Qwen Image 2512 is an improved version of Qwen Image with more realistic human generation, finer textures, and stronger text rendering

    text-to-image

    Modalities: Text, Image

    $0.024/image

  • Qwen Image 3

    Alibaba (Qwen)

    New

    Qwen-Image-3.0 generates and edits images with accurate text rendering, complex layouts, and photographic detail.

    text-to-image

    Modalities: Text, Image

    $0.036/image

  • Qwen Image 3 Pro

    Alibaba (Qwen)

    New

    Qwen-Image-3.0-Pro generates and edits images with dense, accurate text rendering, complex multi-element layouts, and photographic detail.

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • Rd Animation

    Community

    New

    Style consistent animated pixel art sprite generation

    retro-diffusiontext-to-image

    Modalities: Text, Image

    $0.084/image

  • New

    Affordable and fast images

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.0264/image

  • Recraft V3

    Recraft

    New

    State-of-the-art image generation optimized for design and branding. SVG vector output support.

    designvectorbranding

    Modalities: Image

    $0.048/image

  • Recraft V4

    Recraft

    New

    Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and cost-efficient at standard resolution.

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.048/image

  • New

    Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4, with higher resolution for print-ready and large-scale work.

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.30/image

  • New

    Recraft's latest image generation model, built around design taste. Strong prompt accuracy, art-directed composition, and integrated text rendering. Fast and cost-efficient at standard resolution.

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.048/image

  • New

    Recraft's latest image generation model at ~2048px resolution. Same design taste and prompt accuracy as V4.1, with higher resolution for print-ready and large-scale work.

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.30/image

  • New

    A faster, lighter Recraft image generation model optimized for high-volume and production pipelines. Same design taste as V4.1, built for speed and throughput.

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.048/image

  • New

    A faster, lighter Recraft image generation model at ~2048px resolution, optimized for high-volume production. Design taste and prompt accuracy at high resolution with better throughput.

    recraft-aitext-to-image

    Modalities: Text, Image

    $0.30/image

  • Seedream 4

    ByteDance

    New

    ByteDance's unified text-to-image generation and editing model with up to 4K resolution.

    text-to-image

    Modalities: Text, Image

    $0.036/image

  • Seedream 4.5

    ByteDance

    New

    ByteDance Seedream 4.5: upgraded image generation and editing with stronger spatial understanding and world knowledge.

    text-to-image

    Modalities: Text, Image

    $0.048/image

  • New

    ByteDance Seedream 5 Lite: fast, affordable image generation and editing at 2K or 3K resolution.

    text-to-image

    Modalities: Text, Image

    $0.042/image

  • Stable Diffusion 3

    Stability AI

    New

    A text-to-image model with greatly improved performance in image quality, typography, complex prompt understanding, and resource-efficiency

    text-to-image

    Modalities: Text, Image

    $0.042/image

  • New

    2.5 billion parameter image model with improved MMDiT-X architecture

    text-to-image

    Modalities: Text, Image

    $0.042/image

  • Wan 2.2 Image

    Community

    New

    This model generates beautiful cinematic 2 megapixel images in 3-4 seconds and is derived from the Wan 2.2 model through optimisation techniques from the pruna package

    prunaaitext-to-image

    Modalities: Text, Image

    $0.024/image

  • Wan 2.7 Image

    Alibaba (Wan)

    New

    Generate and edit images with Alibaba's Wan 2.7

    wan-videotext-to-image

    Modalities: Text, Image

    $0.036/image

  • Wan 2.7 Image Pro

    Alibaba (Wan)

    New

    Generate and edit high-quality images with Alibaba's Wan 2.7 Pro with 4K output, thinking mode, text-to-image, multi-image editing, and image set generation

    wan-videotext-to-image

    Modalities: Text, Image

    $0.036/image

  • Background removal model from 851-Labs that outputs a clean cutout with a transparent alpha channel. One of the most-run background removers on Replicate, handles people, products and objects on busy backgrounds.

    background-removal851-labscutout

    Modalities: Image

    β‰ˆ $0.00050/run

  • Product advertising photo generator. You upload a cut-out product shot and a prompt describing the scene; it places the product on a new generated background with matching lighting and shadows, so a plain packshot becomes an ecommerce or ad-ready hero image without a photo studio.

    productecommerceproduct-photo

    Modalities: Text, Image

    β‰ˆ $0.0672/run

  • AuraFlow v0.3

    Community

    fal.ai's fully open-source 6.8B flow-based text-to-image model. Up to 1536x1536 resolution.

    auraflowtext-to-imageopen-weights

    Modalities: Text, Image

    β‰ˆ $0.0648/run

  • BiRefNet high-resolution dichotomous image segmentation for background removal. Bilateral reference network that produces sharp matting on fine detail like hair, fur and thin structures, often cleaner than older U2Net or rembg models.

    background-removalbirefnetsegmentation

    Modalities: Image

    β‰ˆ $0.0036/run

  • BRIA AI's commercial background removal model trained on fully licensed data. Produces accurate cutouts for e-commerce and design, with attention to clean edges around products and people.

    background-removalecommercecutout

    Modalities: Image

    $0.0216/image

  • Microsoft Research pipeline by Ziyu Wan et al. that restores scanned old photos, removing scratches, dust and fading and optionally enhancing faces in one pass.

    restoreold-photoscratch-removal

    Modalities: Image

    β‰ˆ $0.010/run

  • Cartoonify

    Community

    catacolabs Cartoonify turns a photo into a flat cartoon illustration. Takes a single image and returns a stylized cartoon version with clean shapes and bold outlines. Straightforward one-input model for avatars and profile pictures.

    avatarportraitcartoon

    Modalities: Image

    β‰ˆ $0.096/run

  • Content-Consistent Super-Resolution model. Reduces hallucination compared to typical diffusion-based upscalers while keeping perceptual quality high.

    upscalingimage-restoreopen-source

    Modalities: Image

    β‰ˆ $0.2761/run

  • High-resolution image upscaler with creative detail re-imagination via SD-based hallucination. Strong for photography and product shots.

    upscalingcreativeopen-source

    Modalities: Image

    β‰ˆ $0.024/run

  • CodeFormer

    Community

    Robust face-restoration model using a transformer-based codebook prior. Handles severe degradation, occlusion, and old-photo restoration with adjustable fidelity-quality tradeoff.

    face-restoreupscalingopen-source

    Modalities: Image

    β‰ˆ $0.0041/run

  • fofr's model generates the same character in many poses and angles from one reference image. Useful for building an avatar set or character sheet where the face and design stay consistent across outputs. Can produce a grid or individual images.

    avatarportraitcharacter

    Modalities: Text, Image

    β‰ˆ $0.0038/run

  • ControlNet conditioned on Canny edge maps. Preserves composition and outlines while restyling with Stable Diffusion 1.5 or SDXL backbones.

    style-transferimage-editcontrolnet

    Modalities: Text, Image

    β‰ˆ $0.0697/run

  • ControlNet conditioned on depth maps. Preserves the 3D scene layout while letting the prompt change style, lighting and content.

    style-transferimage-editcontrolnet

    Modalities: Text, Image

    β‰ˆ $0.0144/run

  • DDColor

    Community

    DDColor by Xiaoyang Kang et al. colorizes black-and-white photos using dual decoders that jointly learn pixel colors and semantic color queries, giving vivid and natural results on old images.

    colorizerestoreddcolor

    Modalities: Image

    β‰ˆ $0.0012/run

  • DreamGaussian

    Community

    Generative Gaussian-splatting model for fast image-to-3D synthesis. Produces textured meshes in two minutes via differentiable rasterization.

    3d-generationimage-to-3dgaussian-splatting

    Modalities: Image

    $0.00168/GPU s

  • Face to Many

    Community

    fofr's face stylizer converts a face photo into 3D render, emoji, pixel art, video-game character, claymation or toy styles. Uses InstantID plus style LoRAs on SDXL to keep the likeness while applying a chosen art style. Popular for fun avatars.

    avatarportraitstylize

    Modalities: Text, Image

    β‰ˆ $0.0105/run

  • fofr's model turns a face photo into a die-cut sticker with a white border and transparent background. Uses InstantID to hold the likeness and outputs a clean PNG suitable for chat stickers or print. Simple single-image input.

    avatarportraitsticker

    Modalities: Text, Image

    β‰ˆ $0.0289/run

  • FLUX PuLID

    Community

    PuLID identity customization running on FLUX.1-dev. Inserts a face from one reference photo into prompt-driven scenes using contrastive alignment, giving higher likeness and detail than SDXL-era ID adapters. Good for realistic avatars and character portraits.

    avatarportraitpulid

    Modalities: Text, Image

    β‰ˆ $0.018/run

  • FLUX.1 [dev]

    Black Forest Labs

    The open-weight 12B rectified-flow transformer from Black Forest Labs. Close to FLUX Pro quality with a guidance-distilled checkpoint released under a non-commercial license. The most widely fine-tuned base in the FLUX family.

    fluxtext-to-imageopen-weights

    Modalities: Text, Image

    $0.030/image

  • FLUX.1 [schnell]

    Black Forest Labs

    The fastest FLUX model from Black Forest Labs, distilled to produce images in 1 to 4 steps. Apache 2.0 licensed for commercial use. Built for high-volume generation and real-time previews.

    fluxtext-to-imagefast

    Modalities: Text, Image

    $0.0036/image

  • FLUX.1 Canny

    Black Forest Labs

    FLUX structural control via Canny edge maps. Preserve composition while restyling.

    fluximage-editcontrolnet

    Modalities: Text, Image

    $0.060/image

  • FLUX.1 Canny [dev]

    Black Forest Labs

    Open-weight edge-guided FLUX model from Black Forest Labs. Extracts Canny edges from a control image and regenerates it from your prompt while holding the original composition and outlines, so you can restyle a scene without changing its structure.

    fluximage-editcontrolnet

    Modalities: Text, Image

    $0.030/image

  • FLUX.1 Depth

    Black Forest Labs

    FLUX structural control via depth maps. Keep 3D scene layout while changing style/content.

    fluximage-editcontrolnet

    Modalities: Text, Image

    $0.060/image

  • FLUX.1 Depth [dev]

    Black Forest Labs

    Open-weight depth-guided FLUX model from Black Forest Labs. Derives a depth map from the control image and regenerates from your prompt while preserving 3D spatial layout, useful for re-texturing rooms, products, or scenes without moving objects.

    fluximage-editcontrolnet

    Modalities: Text, Image

    $0.030/image

  • FLUX.1 Fill

    Black Forest Labs

    Black Forest Labs' inpainting/outpainting model for FLUX. Fill masked regions with prompt-guided content.

    fluximage-editinpainting

    Modalities: Text, Image

    $0.060/image

  • FLUX.1 Fill [dev]

    Black Forest Labs

    Black Forest Labs' open-weight inpainting and outpainting model, guidance-distilled from FLUX.1 Fill [pro]. You supply an image plus a mask and a prompt; it fills the masked region or extends the canvas with prompt-guided content that matches lighting and texture.

    fluximage-editinpainting

    Modalities: Text, Image

    $0.048/image

  • FLUX.1 Kontext [dev]

    Black Forest Labs

    Open-weight version of FLUX.1 Kontext by Black Forest Labs. Instruction-based editing: pass an input image and a plain text edit ('change the jacket to red', 'remove the person on the left') and it applies the change while keeping the rest of the scene and identity consistent.

    fluxkontextimage-edit

    Modalities: Text, Image

    $0.030/image

  • FLUX.1 Redux

    Black Forest Labs

    FLUX image-variation adapter. Generate variations and remixes from a reference image.

    fluximage-editvariation

    Modalities: Text, Image

    $0.0036/image

  • FLUX.1-dev inpainting wrapper that fills masked parts of an image from a prompt. Useful when you want FLUX-quality fills with a simple image plus mask plus prompt interface and adjustable mask strength.

    fluximage-editinpainting

    Modalities: Text, Image

    β‰ˆ $0.090/run

  • GFPGAN v1.4

    Tencent ARC

    Tencent ARC face-restoration GAN. Reconstructs realistic facial detail in low-quality or compressed photos using a pretrained StyleGAN2 prior.

    face-restoreupscalingopen-source

    Modalities: Image

    β‰ˆ $0.0027/run

  • Tencent's Hunyuan3D 2.0 image-to-3D pipeline. Two-stage shape and texture generation producing high-resolution textured meshes.

    3d-generationimage-to-3dopen-weights

    Modalities: Image

    β‰ˆ $0.156/run

  • Lvmin Zhang's IC-Light packaged by zsxkib. Relights a product or portrait from a text prompt or a chosen light direction while keeping the subject's shape and detail intact, so a flat product photo can be given studio, window, or dramatic side lighting without re-shooting.

    productic-lightrelight

    Modalities: Text, Image

    β‰ˆ $0.0781/run

  • Ideogram v2

    Ideogram

    Ideogram's text-to-image model known for accurate in-image text and typography. Handles posters, logos, and signage where other models garble lettering. Supports magic prompt expansion and multiple aspect ratios.

    text-to-imagetypographylogos

    Modalities: Text, Image

    $0.096/image

  • Ideogram's fast v3 model, the fastest and cheapest tier of the v3 family. Known for accurate in-image text rendering and reliable typography, which most diffusion models still get wrong. Hosted API only.

    text-to-imagetypographytext-in-image

    Modalities: Image

    $0.036/image

  • IDM-VTON virtual try-on from the CVPR 2024 paper. You give it a photo of a person and a garment image; it dresses the person in that garment while preserving pose, body shape, and the garment's pattern and text. Good for showing a clothing product on a model for an ecommerce listing.

    productvirtual-try-onvton

    Modalities: Text, Image

    β‰ˆ $0.0277/run

  • Berkeley InstructPix2Pix. Edits an image from natural-language instructions in a single forward pass. Trained on GPT-3 plus Stable Diffusion synthetic pairs.

    style-transferimage-editinstruction-edit

    Modalities: Text, Image

    β‰ˆ $0.0039/run

  • Tencent's face-identity conditioning adapter for SD/SDXL. Face embedding + CLIP for ID-consistent generation.

    tencentimage-editface-id

    Modalities: Text, Image

    β‰ˆ $0.0277/run

  • Janus Pro 7B

    DeepSeek

    DeepSeek's unified multimodal model. Decouples vision encoding for both understanding and generation tasks.

    janusopen-weightsunified-multimodal

    Modalities: Text, Image

    β‰ˆ $0.0169/run

  • Kuaishou's bilingual (CN/EN) latent diffusion text-to-image model with strong text rendering.

    kuaishoutext-to-imageopen-weights

    Modalities: Text, Image

    β‰ˆ $0.0889/run

  • SDXL fine-tune by mejiabrayan aimed at logo generation. Produces simple, centered mark and wordmark style logos from a text prompt. Useful for quick brand concepts and mockups. Raster PNG output, not vector.

    logoiconsdxl

    Modalities: Image

    β‰ˆ $0.0372/run

  • Detail-hallucinating upscaler in the Magnific style. Adds plausible high-frequency texture using a Stable Diffusion refiner conditioned on the low-res input.

    upscalingcreativemagnific-style

    Modalities: Image

    β‰ˆ $0.0589/run

  • OOTDiffusion virtual try-on. Takes a clear photo of a model and an upper-body garment and renders the garment onto the person using an outfitting-fusion diffusion approach that keeps the garment's texture and the model's pose. A lightweight alternative to IDM-VTON for clothing previews.

    productvirtual-try-onvton

    Modalities: Image

    β‰ˆ $0.1801/run

  • PhotoMaker

    Tencent ARC

    Tencent ARC PhotoMaker. Identity-preserving stylized photo generation from a stacked-ID embedding. Realistic re-styling of a subject in seconds.

    style-transferimage-editidentity-preserving

    Modalities: Text, Image

    β‰ˆ $0.0080/run

  • Playground AI's diffusion model tuned for aesthetics. SDXL-based architecture trained on the EDM formulation, rated by users as more visually pleasing than SDXL in their study. Strong on vivid color and contrast.

    text-to-imageaestheticopen-weights

    Modalities: Image

    β‰ˆ $0.0732/run

  • Point-E

    Community

    OpenAI Point-E text-to-point-cloud system. Fast 3D point-cloud generation from text, optionally lifted to a mesh via marching cubes.

    3d-generationopenaipoint-cloud

    Modalities: Text, Image

    β‰ˆ $0.0708/run

  • Qwen-Image-Edit

    Alibaba (Qwen)

    Alibaba Qwen's instruction-driven image editor. Extends Qwen-Image's text-rendering ability to editing, so it handles both semantic edits (swap objects, change style) and precise text edits inside the image while preserving the original layout and unedited regions.

    image-editinstructtext-editing

    Modalities: Text, Image

    $0.036/image

  • AI-Upscaler that increases image resolution up to 4x while preserving texture and detail. Trained on synthetic and real data to reduce common ESRGAN artifacts.

    upscalingimage-restoreopen-source

    Modalities: Image

    $0.0024/image

  • Real-ESRGAN variant fine-tuned for anime, manga, and illustrated artwork. 4x upscaling with cartoon-aware artifact suppression.

    upscalinganimeopen-source

    Modalities: Image

    β‰ˆ $0.0051/run

  • Recraft's v3 variant that outputs vector SVG instead of raster pixels. Generates clean, editable logos, icons and illustrations that scale without quality loss, which is unusual among image models. Hosted API only.

    text-to-imagesvgvector

    Modalities: Image

    $0.096/image

  • Recraft V4 SVG turns a text prompt into production-ready SVG vector art with clean geometry and structured, editable layers. Newer generation than V3 with improved design quality on logos, icons and flat illustration. Returns true vector paths, not a traced bitmap.

    svgvectortext-to-image

    Modalities: Text, Image

    $0.096/image

  • Rembg

    Community

    Open-source background-removal tool wrapping U2Net. Produces alpha mattes for photos, products and people with no manual masking.

    background-removalmattingopen-source

    Modalities: Image

    β‰ˆ $0.0047/run

  • Lucataco's remove-bg, a rembg-based background removal model that returns the foreground subject on a transparent background. A popular, low-cost option for quick product and portrait cutouts.

    background-removallucatacorembg

    Modalities: Image

    β‰ˆ $0.00060/run

  • Object removal and cleanup using LaMa inpainting. Paint a mask over an unwanted object, logo or person and the model fills the area with plausible background, erasing it from the photo.

    background-removalobject-removallama

    Modalities: Image

    β‰ˆ $0.00090/run

  • SDXL Emoji

    Community

    SDXL fine-tune by fofr trained on Apple emoji art. Generates rounded, glossy emoji and icon style graphics from a text prompt, useful for custom reactions, app glyphs and playful icon sets. Raster output.

    emojiiconsdxl

    Modalities: Image

    β‰ˆ $0.0156/run

  • SDXL inpainting built on the Hugging Face Diffusers inpaint pipeline. Replace or remove masked regions of an image with prompt-conditioned content at SDXL resolution. A cheap, well-understood baseline for object removal and local edits.

    sdxlstability-aiimage-edit

    Modalities: Text, Image

    β‰ˆ $0.0026/run

  • OpenAI Shap-E text/image to 3D. Generates implicit neural representations renderable as textured meshes or NeRFs.

    3d-generationopenaiopen-source

    Modalities: Text, Image

    β‰ˆ $0.1116/run

  • Stability AI's 8B MMDiT-based flagship. Open weights at 1MP with improved typography and prompt adherence over SDXL. The largest model in the SD 3.5 release line.

    stable-diffusiontext-to-imageopen-weights

    Modalities: Text, Image

    $0.078/image

  • Distilled, 4-step version of SD 3.5 Large from Stability AI. Keeps most of the large model's quality and text rendering at a fraction of the inference time. Open weights under the Stability Community License.

    stable-diffusiontext-to-imagefast

    Modalities: Text, Image

    $0.048/image

  • Stability AI's SDXL model via Replicate. High-quality image generation with extensive customization.

    open-sourcecustomizable

    Modalities: Image

    β‰ˆ $0.0015/run

  • StarVector 8B is a multimodal model that generates SVG code directly from an input image. Rather than tracing pixels, it predicts the SVG markup token by token, which can produce compact, semantically structured paths for icons and simple graphics. Research model from the StarVector project.

    svgvectorstarvector

    Modalities: Image

    $0.00117/GPU s

  • SUPIR

    Community

    SUPIR by Fanghua Yu et al. is a large diffusion-based restoration model that recovers photorealistic detail from heavily degraded images and can be steered with a text prompt describing the scene.

    restoresuper-resolutionsupir

    Modalities: Text, Image

    β‰ˆ $0.4921/run

  • SUPIR (Scaling-Up Image Restoration) photo-real restoration model. Combines SDXL prior with language-guided controls for severely degraded inputs.

    upscalingimage-restoreopen-source

    Modalities: Image

    β‰ˆ $0.30/run

  • Swin2SR

    Community

    Transformer-based image super-resolution using Swin-V2 attention. Handles classical, lightweight, real-world, and compressed-input variants with 2x/4x upscaling.

    upscalingtransformeropen-source

    Modalities: Image

    β‰ˆ $0.0088/run

  • TRELLIS (3D)

    Community

    Microsoft TRELLIS image-to-3D model. Generates textured 3D assets in GLB or Gaussian-splat format from a single reference image.

    3d-generationimage-to-3dopen-source

    Modalities: Image

    β‰ˆ $0.042/run

  • PNG/JPG to SVG vectorizer built on VTracer, the open-source raster-to-vector engine. Traces a bitmap into layered color regions and clean paths with controls for color count, area threshold and path simplification. Fast, deterministic alternative to model-based vectorizers.

    svgvectorvtracer

    Modalities: Image

    β‰ˆ $0.0022/run

28 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Image generation models for product, marketing, and design

Image models turn a text prompt β€” and optionally a reference image or a mask β€” into a finished raster. The category covers everything from photoreal product shots to vector-style illustrations to controllable inpainting and outpainting. You reach for an image model when you need on-brand visuals at scale, when a designer's queue is the bottleneck, or when you want to ship a generative feature inside your own product.

Unlike text models, image generation is billed per-call rather than per-token. On Railwail a single image currently costs from under half a cent (FLUX Schnell, $0.0036) to about eleven cents (Ideogram v3 Quality, $0.108); FLUX 1.1 Pro and Imagen 4 cost $0.048 per image. Higher resolutions and longer step counts cost proportionally more. Some providers expose a separate edit endpoint at a different rate; check the model card before integrating.

The core trade-off is photorealism versus controllability. Diffusion flagships (FLUX 1.1 Pro, Imagen 4, Recraft V3) produce magazine-quality output but ignore detailed compositional instructions about half the time. Smaller models (SDXL, Playground V3, Stable Diffusion 3.5) cost ten times less, render in under two seconds, and let you steer the result with ControlNet, IP-Adapter, or LoRA. For batch production work, the smaller and steerable pipeline almost always wins; for one-off hero shots, reach for the flagship.

Watch out for context dilution in image prompts: most diffusion models cap useful prompt length around 75 tokens, so jamming in twelve adjectives and three style references typically averages them all out instead of stacking them. Write the subject, the action, and the lighting first; everything after the third clause has diminishing influence on the result.

Licensing matters: most providers grant a perpetual commercial-use license on generated images, but a few (FLUX Schnell free tier, some open checkpoints) restrict to non-commercial. The model card spells it out β€” read it before you put output on a billboard.

Top picks below cover the photorealism flagship, the cheapest workhorse, the longest-prompt model, and the fastest realtime option in the category.

Typical tasks

  • On-brand product photography
  • Marketing and ad creative at scale
  • Editorial illustration
  • App and game asset generation
  • AI-powered avatar and character art
  • Storyboard and concept-art workflows

Model comparisons

Frequently asked questions

Which image model is the fastest?

Turbo and lightning variants β€” SDXL Turbo, FLUX Schnell, Ideogram Turbo β€” render a 1024px image in 1-3 seconds. Standard flagship pipelines (FLUX Pro, Imagen 4) take 6-15 seconds. Pick the turbo tier when you're previewing in a UI and the flagship when the image is final.

Which model is best for photorealism?

FLUX 1.1 Pro and Imagen 4 are strong on skin texture, lighting realism, and hand anatomy. Recraft V3 wins on product photography and white-background commerce shots. For full photoreal portraits, FLUX 1.1 Pro Ultra adds another step of detail at roughly twice the cost.

Is pricing per-image or per-token?

Image generation is billed per-call. A standard 1024Γ—1024 image is one call. Higher resolutions, more diffusion steps, and edit operations may multiply the price β€” check the model card for the exact rate.

What resolutions are supported?

Standard tiers ship 1024Γ—1024 and common aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4). Pro tiers add 2K and 4K outputs and arbitrary aspect ratios up to 2048 on the long edge. Upscalers in the image-edit category push results to 4K and beyond.

Can I edit existing images?

Yes β€” inpainting, outpainting, controlled edits via mask + prompt, and reference-image conditioning are available with the right model. The /models/image page includes models tuned for editing, such as FLUX Fill, FLUX Canny, FLUX Depth, and IP-Adapter variants.

Can I control the style?

Three ways: (1) prompt-engineering with style references ('in the style of...'), (2) LoRAs or style adapters where supported, and (3) IP-Adapter and FLUX Redux for image-to-image style transfer. Recraft V3 and Ideogram offer named style presets directly in the API.

Can I use generated images commercially?

Usually yes, but it depends on the model: most commercial image APIs allow commercial use of the output, while some open-weights checkpoints (for example research previews) are limited to non-commercial use. Check the license on the provider's model page before you ship.

Is there a speed vs quality knob I can tune?

Yes: most diffusion models expose a `num_inference_steps` (or equivalent) parameter. Default is usually 20-30 steps for a good balance. Drop to 4-8 steps with a turbo variant for instant previews; bump to 50+ steps on a Pro tier for final renders. Quality plateaus around step 50 for most models.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.