Multimodal Models

Models that read images, documents or video alongside text.

Available
38of 69
Providers
6
Price range
$0.00030 – $0.1021per default run

38 models

  • New

    Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    reasoningagenticvision

    Modalities: Text, Image1M context

    $6.00/1M in

    $30.00/1M out

  • New

    Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    balancedproductionagentic

    Modalities: Text, Image1M context

    $3.60/1M in

    $18.00/1M out

  • Gemini 3 Flash

    Google DeepMind

    New

    Google's April 2026 fast multimodal model. Combines Gemini 3 Pro's reasoning with Flash-tier latency and price. Default model in the Gemini app.

    balancedmultimodallow-latency

    Modalities: Text, Image, Video, Audio1M context

    $0.60/1M in

    $3.60/1M out

  • Gemini 3.1 Pro

    Google DeepMind

    New

    Google DeepMind's February 2026 flagship. 1M-token (1,048,576) context, native multimodal (text/image/audio/video), Deep Think reasoning.

    multimodaldeep-thinklong-context

    Modalities: Text, Image, Video, Audio1M context

    $2.40/1M in

    $14.40/1M out

  • GPT-5.4

    OpenAI

    New

    OpenAI's unified flagship combining GPT and o-series reasoning into one model. 1M context, multimodal, top SWE-Bench Pro and OSWorld scores.

    reasoningagenticvision

    Modalities: Text, Image1.1M context

    $3.00/1M in

    $18.00/1M out

  • New

    OpenAI's efficient mid-tier model. 2x faster than its predecessor, 400k context, approaches GPT-5.4 quality on SWE-Bench Pro at a fraction of the cost.

    balancedcost-efficientvision

    Modalities: Text, Image400K context

    $0.90/1M in

    $5.40/1M out

  • BLIP

    Salesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    blipcaptioningvqa

    Modalities: Text, Image

    β‰ˆ $0.00030/run

  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    clip-interrogatorcaptioningtagging

    Modalities: Text, Image

    β‰ˆ $0.0457/run

  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    depthvision-understandingopen-weights

    Modalities: Text, Image

    β‰ˆ $0.0050/run

  • Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.

    segmentationvision-understandingopen-weights

    Modalities: Text, Image, Video

    β‰ˆ $0.018/run

  • New

    Anthropic's fastest and cheapest 4.x model. Strong vision and tool use at ultra-low latency, ideal for high-concurrency workloads.

    cost-efficientlow-latencyvision

    Modalities: Text, Image200K context

    $1.20/1M in

    $6.00/1M out

  • New

    OpenAI's smallest and cheapest GPT-5.4 variant. Built for high-volume classification, extraction and coding subagents at edge-grade latency.

    cost-efficientlow-latencyvision

    Modalities: Text, Image400K context

    $0.24/1M in

    $1.50/1M out

  • CogVLM2 19B

    Community

    Tsinghua CogVLM2 19B with Llama-3 8B base plus 11B vision expert. Strong document understanding and visual reasoning, 8k context.

    multimodalvision-understandingtsinghua

    Modalities: Text, Image8.2K context

    β‰ˆ $0.0114/run

  • DeepSeek-VL 7B chat model. Vision-language model with hybrid vision encoder and strong real-world visual question answering performance.

    multimodalvision-understandingopen-weights

    Modalities: Text, Image4.1K context

    β‰ˆ $0.0086/run

  • Naver CLOVA Donut OCR-free document-understanding transformer. End-to-end JSON extraction from forms, receipts and invoices without explicit OCR.

    ocrvision-understandingnaver

    Modalities: Text, Image

    β‰ˆ $0.0648/run

  • Dots OCR

    Community

    Rednote Hilab Dots OCR. End-to-end document parsing model with layout, text and reading-order prediction in one transformer.

    ocrvision-understandingopen-source

    Modalities: Text, Image

    β‰ˆ $0.0132/run

  • EasyOCR

    Community

    JaidedAI EasyOCR. Simple Python OCR wrapper supporting 80+ languages with deep-learning text detection and recognition.

    ocrvision-understandingmultilingual

    Modalities: Text, Image

    β‰ˆ $0.0014/run

  • Microsoft Florence-2 Large. Unified prompt-based vision foundation model for captioning, detection, segmentation and OCR with a single 770M-param backbone.

    multimodalvision-understandingmicrosoft

    Modalities: Text, Image

    β‰ˆ $0.0012/run

  • GLPN Depth

    Community

    Global-Local Path Networks depth-estimation model. Combines hierarchical transformer encoder with selective feature fusion for sharp boundaries.

    depthvision-understandingopen-source

    Modalities: Text, Image

    β‰ˆ $0.0094/run

  • GOT-OCR 2.0

    Community

    StepFun GOT-OCR 2.0. Unified end-to-end OCR-2.0 model handling text, formulas, charts, sheet music and geometric shapes in one architecture.

    ocrvision-understandingstepfun

    Modalities: Text, Image

    β‰ˆ $0.0012/run

  • Grounded-SAM

    Community

    Grounding DINO plus SAM. Open-vocabulary text-prompted detection and segmentation in one pipeline for fully-automatic mask generation.

    segmentationvision-understandingopen-vocabulary

    Modalities: Text, Image

    β‰ˆ $0.0015/run

  • Idefics3 8B

    Community

    Hugging Face Idefics3 8B. Llama-3 based open-source vision-language model with strong document QA and chart-understanding performance.

    multimodalvision-understandingopen-weights

    Modalities: Text, Image8.2K context

    β‰ˆ $0.0012/run

  • Meta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.

    metallamavision-understanding

    Modalities: Text, Image

    β‰ˆ $0.0039/run

  • Meta Llama 3.2 90B Vision. Largest open-weights Llama vision model. Strong visual reasoning, chart, OCR and document understanding.

    multimodalvision-understandingmeta

    Modalities: Text, Image131.1K context

    β‰ˆ $0.0070/run

  • LLaVA 1.6 (LLaVA-NeXT) with a Vicuna-13B language backbone. Open vision-language chat model that describes images, answers questions, reads charts and reasons about scenes. Version 1.6 adds higher input resolution and better OCR and reasoning than LLaVA 1.5.

    llavacaptioningvqa

    Modalities: Text, Image4.1K context

    β‰ˆ $0.1021/run

  • Lotus-G

    Community

    Lotus generative depth model. Treats depth as a generation task using a diffusion model, producing higher-fidelity depth on textured surfaces.

    depthvision-understandingdiffusion

    Modalities: Text, Image

    β‰ˆ $0.0408/run

  • Marigold

    Community

    ETH Zurich Marigold. Diffusion-based monocular depth-estimation model fine-tuned from Stable Diffusion with strong fine-detail recovery.

    depthvision-understandingdiffusion

    Modalities: Text, Image

    β‰ˆ $0.0936/run

  • Meta Mask2Former universal image-segmentation transformer. Single architecture for panoptic, instance and semantic segmentation tasks.

    segmentationvision-understandingtransformer

    Modalities: Text, Image

    β‰ˆ $0.0276/run

  • MiDaS v3.1

    Community

    Intel MiDaS v3.1 relative depth-estimation model. Robust zero-shot single-image depth across diverse domains and resolutions.

    depthvision-understandingintel

    Modalities: Text, Image

    β‰ˆ $0.00030/run

  • MiniCPM-V 2.6

    Community

    OpenBMB MiniCPM-V 2.6. 8B vision-language model with strong single-image, multi-image and video understanding plus OCR capabilities.

    multimodalvision-understandingopen-weights

    Modalities: Text, Image, Video32.8K context

    β‰ˆ $0.0015/run

  • Molmo 7B

    Community

    Allen AI Molmo 7B-D on Replicate. Open vision-language model trained on the PixMo data, notable for pointing at and locating objects in images, not just describing them.

    allenaimolmovision-understanding

    Modalities: Text, Image

    β‰ˆ $0.0505/run

  • Moondream2

    Community

    Moondream2 small vision-language model on Replicate. About 1.9B params, designed to run on edge devices, handles captioning, visual QA and short OCR-style reads at very low cost.

    moondreamvision-understandingopen-source

    Modalities: Text, Image

    β‰ˆ $0.0020/run

  • olmOCR

    Community

    Allen AI olmOCR. Open-source 7B vision-language model fine-tuned for high-fidelity document parsing including math, code and tables.

    ocrvision-understandingallenai

    Modalities: Text, Image

    β‰ˆ $0.0205/run

  • OpenPose

    Community

    CMU OpenPose multi-person 2D pose estimator. Real-time keypoint detection for body, hand, face and foot using Part Affinity Fields.

    posevision-understandingcmu

    Modalities: Text, Image, Video

    β‰ˆ $0.0012/run

  • PaddleOCR v3

    Community

    Baidu PaddleOCR v3 PP-OCR pipeline. Lightweight detector plus recognizer optimized for production use with 80+ language support.

    ocrvision-understandingbaidu

    Modalities: Text, Image

    β‰ˆ $0.0216/run

  • Alibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.

    qwenalibabavision-understanding

    Modalities: Text, Image, Video

    β‰ˆ $0.0024/run

  • Segformer B5

    Community

    NVIDIA SegFormer-B5 semantic segmentation. Hierarchical transformer encoder with lightweight MLP decoder, strong ADE20k and Cityscapes results.

    segmentationvision-understandingnvidia

    Modalities: Text, Image

    β‰ˆ $0.00030/run

  • ZoeDepth

    Community

    Intel ZoeDepth metric depth-estimation model. Combines relative-depth pretraining with metric fine-tuning for absolute distance in real units.

    depthvision-understandingintel

    Modalities: Text, Image

    β‰ˆ $0.00060/run

29 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Multimodal models for vision, OCR, and document understanding

Multimodal models accept text plus images (sometimes plus audio or video) and produce text output. Reach for one when your input contains images and your output is structured information: extract an invoice, describe a chart, transcribe a handwritten note, answer questions about a UI screenshot.

Pricing, trade-offs and pitfalls

Pricing is per-token like regular LLMs, with one twist: every image is counted as a fixed number of tokens β€” typically 250-1,500 tokens depending on resolution and detail mode. A standard 1024Γ—1024 image costs roughly the same as a 1,000-word text input. High-detail mode (preserving fine text and small UI elements) costs 2-4Γ— more. Plan budgets accordingly β€” a workload that processes thousands of receipt scans per day adds up quickly at flagship rates; multiply the per-token price on the model page by your image token count.

The trade-off is OCR accuracy, reasoning quality, and cost. Flagships (GPT-5 Vision, Claude 4.6, Gemini 2.5) read complex layouts and reason over chart contents very reliably. Specialized OCR-first models (Qwen 2.5 VL, Pixtral, InternVL) sometimes outperform on pure text extraction at a fraction of the cost. For pure document-to-JSON pipelines, a specialized model with a strict JSON schema usually wins. For document-and-reasoning workloads ('extract this invoice AND tell me if the tax math is right'), flagships win.

Watch out for resolution limits: most models downscale very large images before processing, which can destroy fine text in screenshots and dense documents. Pre-process β€” slice tall documents into single-page images, upscale low-resolution scans before sending β€” to preserve readability. Also watch out for hallucinations on hard-to-read regions; multimodal models tend to confidently invent text where the source is illegible.

Top picks above cover the most accurate flagship, the cheapest workhorse, the highest-resolution supporter, and the fastest streaming option.

Typical tasks

  • Invoice and receipt extraction
  • Document OCR and structuring
  • UI screenshot analysis
  • Chart and graph reading
  • Medical and scientific image annotation
  • Visual QA for accessibility

Frequently asked questions

Which multimodal model reads documents best?

GPT-5 Vision and Claude 4.6 lead on complex layouts, charts, and handwriting. For pure OCR on clean documents, Qwen 2.5 VL and Pixtral often match flagship quality at one-tenth the cost. Run a sample of your own documents on each before committing.

How is pricing calculated for images?

Every image is converted to a fixed token count β€” typically 250-1,500 tokens depending on resolution and detail mode. So a 1024Γ—1024 image costs the same as ~1,000 tokens of text. Multi-page documents multiply linearly; check each model card for exact image-token rates.

What image formats are accepted?

JPEG, PNG, WebP, and GIF (first frame) are universal. Some models also accept PDF directly (one image per page). Max file size is typically 20 MB; max dimensions 8K on a side. Pre-resize larger files to avoid silent downscaling.

Can it read handwriting?

Flagships handle clear handwriting reliably and most cursive scripts in major languages. Quality drops on heavily stylized writing, faded ink, and complex shorthand. For high-volume handwriting workloads, fine-tune on a sample of your domain or pair with a specialized OCR pre-processor.

Can it understand charts and graphs?

Yes β€” chart reading is one of the strongest use-cases. Flagships extract data points from bar charts, line graphs, and pie charts; they reason about axis scales and identify trends. For dense scientific figures, ask for structured output (CSV or JSON) rather than narrative description.

How does it handle low-quality scans?

Quality degrades sharply below 150 DPI or above 30% noise. Pre-process with deskewing, contrast normalization, and lightweight denoising. For very poor scans, run an upscaler first (from /models/image) to restore readability before sending.

Can I send multiple images in one request?

Yes β€” most flagships accept 10-20 images in a single prompt. Useful for comparing screenshots, processing multi-page documents in one pass, or analyzing sequences of frames. Total token count (including all image tokens) must fit inside the context window.

Is video frame-by-frame supported?

Some multimodal models (Gemini 2.5 Pro, Qwen 2.5 VL) accept video directly and sample frames internally. For others, extract frames yourself (every 1-2 seconds is typical) and send them as a multi-image prompt. For real-time video analysis, look at dedicated video-understanding models.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.