Claude Opus 4.7
Anthropic
Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.
reasoningagenticvision
$6.00/1M in
$30.00/1M out
Models that read images, documents or video alongside text.
Quick picks
38 models
Anthropic
Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.
reasoningagenticvision
$6.00/1M in
$30.00/1M out
Anthropic
Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.
balancedproductionagentic
$3.60/1M in
$18.00/1M out
Google DeepMind
Google's April 2026 fast multimodal model. Combines Gemini 3 Pro's reasoning with Flash-tier latency and price. Default model in the Gemini app.
balancedmultimodallow-latency
$0.60/1M in
$3.60/1M out
Google DeepMind
Google DeepMind's February 2026 flagship. 1M-token (1,048,576) context, native multimodal (text/image/audio/video), Deep Think reasoning.
multimodaldeep-thinklong-context
$2.40/1M in
$14.40/1M out
OpenAI
OpenAI's unified flagship combining GPT and o-series reasoning into one model. 1M context, multimodal, top SWE-Bench Pro and OSWorld scores.
reasoningagenticvision
$3.00/1M in
$18.00/1M out
OpenAI
OpenAI's efficient mid-tier model. 2x faster than its predecessor, 400k context, approaches GPT-5.4 quality on SWE-Bench Pro at a fraction of the cost.
balancedcost-efficientvision
$0.90/1M in
$5.40/1M out
Salesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
blipcaptioningvqa
β $0.00030/run
Community
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
clip-interrogatorcaptioningtagging
β $0.0457/run
Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
depthvision-understandingopen-weights
β $0.0050/run
Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.
segmentationvision-understandingopen-weights
β $0.018/run
Anthropic
Anthropic's fastest and cheapest 4.x model. Strong vision and tool use at ultra-low latency, ideal for high-concurrency workloads.
cost-efficientlow-latencyvision
$1.20/1M in
$6.00/1M out
OpenAI
OpenAI's smallest and cheapest GPT-5.4 variant. Built for high-volume classification, extraction and coding subagents at edge-grade latency.
cost-efficientlow-latencyvision
$0.24/1M in
$1.50/1M out
Community
Tsinghua CogVLM2 19B with Llama-3 8B base plus 11B vision expert. Strong document understanding and visual reasoning, 8k context.
multimodalvision-understandingtsinghua
β $0.0114/run
DeepSeek
DeepSeek-VL 7B chat model. Vision-language model with hybrid vision encoder and strong real-world visual question answering performance.
multimodalvision-understandingopen-weights
β $0.0086/run
Community
Naver CLOVA Donut OCR-free document-understanding transformer. End-to-end JSON extraction from forms, receipts and invoices without explicit OCR.
ocrvision-understandingnaver
β $0.0648/run
Community
Rednote Hilab Dots OCR. End-to-end document parsing model with layout, text and reading-order prediction in one transformer.
ocrvision-understandingopen-source
β $0.0132/run
Community
JaidedAI EasyOCR. Simple Python OCR wrapper supporting 80+ languages with deep-learning text detection and recognition.
ocrvision-understandingmultilingual
β $0.0014/run
Community
Microsoft Florence-2 Large. Unified prompt-based vision foundation model for captioning, detection, segmentation and OCR with a single 770M-param backbone.
multimodalvision-understandingmicrosoft
β $0.0012/run
Community
Global-Local Path Networks depth-estimation model. Combines hierarchical transformer encoder with selective feature fusion for sharp boundaries.
depthvision-understandingopen-source
β $0.0094/run
Community
StepFun GOT-OCR 2.0. Unified end-to-end OCR-2.0 model handling text, formulas, charts, sheet music and geometric shapes in one architecture.
ocrvision-understandingstepfun
β $0.0012/run
Community
Grounding DINO plus SAM. Open-vocabulary text-prompted detection and segmentation in one pipeline for fully-automatic mask generation.
segmentationvision-understandingopen-vocabulary
β $0.0015/run
Community
Hugging Face Idefics3 8B. Llama-3 based open-source vision-language model with strong document QA and chart-understanding performance.
multimodalvision-understandingopen-weights
β $0.0012/run
Community
Meta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.
metallamavision-understanding
β $0.0039/run
Community
Meta Llama 3.2 90B Vision. Largest open-weights Llama vision model. Strong visual reasoning, chart, OCR and document understanding.
multimodalvision-understandingmeta
β $0.0070/run
Community
LLaVA 1.6 (LLaVA-NeXT) with a Vicuna-13B language backbone. Open vision-language chat model that describes images, answers questions, reads charts and reasons about scenes. Version 1.6 adds higher input resolution and better OCR and reasoning than LLaVA 1.5.
llavacaptioningvqa
β $0.1021/run
Community
Lotus generative depth model. Treats depth as a generation task using a diffusion model, producing higher-fidelity depth on textured surfaces.
depthvision-understandingdiffusion
β $0.0408/run
Community
ETH Zurich Marigold. Diffusion-based monocular depth-estimation model fine-tuned from Stable Diffusion with strong fine-detail recovery.
depthvision-understandingdiffusion
β $0.0936/run
Meta
Meta Mask2Former universal image-segmentation transformer. Single architecture for panoptic, instance and semantic segmentation tasks.
segmentationvision-understandingtransformer
β $0.0276/run
Community
Intel MiDaS v3.1 relative depth-estimation model. Robust zero-shot single-image depth across diverse domains and resolutions.
depthvision-understandingintel
β $0.00030/run
Community
OpenBMB MiniCPM-V 2.6. 8B vision-language model with strong single-image, multi-image and video understanding plus OCR capabilities.
multimodalvision-understandingopen-weights
β $0.0015/run
Community
Allen AI Molmo 7B-D on Replicate. Open vision-language model trained on the PixMo data, notable for pointing at and locating objects in images, not just describing them.
allenaimolmovision-understanding
β $0.0505/run
Community
Moondream2 small vision-language model on Replicate. About 1.9B params, designed to run on edge devices, handles captioning, visual QA and short OCR-style reads at very low cost.
moondreamvision-understandingopen-source
β $0.0020/run
Community
Allen AI olmOCR. Open-source 7B vision-language model fine-tuned for high-fidelity document parsing including math, code and tables.
ocrvision-understandingallenai
β $0.0205/run
Community
CMU OpenPose multi-person 2D pose estimator. Real-time keypoint detection for body, hand, face and foot using Part Affinity Fields.
posevision-understandingcmu
β $0.0012/run
Community
Baidu PaddleOCR v3 PP-OCR pipeline. Lightweight detector plus recognizer optimized for production use with 80+ language support.
ocrvision-understandingbaidu
β $0.0216/run
Community
Alibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.
qwenalibabavision-understanding
β $0.0024/run
Community
NVIDIA SegFormer-B5 semantic segmentation. Hierarchical transformer encoder with lightweight MLP decoder, strong ADE20k and Cityscapes results.
segmentationvision-understandingnvidia
β $0.00030/run
Community
Intel ZoeDepth metric depth-estimation model. Combines relative-depth pretraining with metric fine-tuning for absolute distance in real units.
depthvision-understandingintel
β $0.00060/run
xAI
xAI's May 2026 flagship. 1M context, vision, always-on reasoning, real-time X/web retrieval via DeepSearch.
reasoningvisiondeep-search
Currently not offered
Anthropic
Anthropic Claude 3.5 Sonnet with image input. 200k context, strong on dense documents, tables, charts and handwriting. Reliable structured extraction from screenshots and scans.
visionmultimodaldocument-qa
Currently not offered
Google DeepMind
Google Gemini 1.5 Pro with native multimodal input. Reads images, long PDFs, audio and video in up to a 2M-token context, useful for whole-document and long-video understanding.
geminivisionmultimodal
Currently not offered
xAI's cost-efficient high-throughput model. 2M context, optional reasoning, optimized for agentic loops and real-time apps.
cost-efficientvisiondeep-search
Deactivated
Salesforce
Salesforce BLIP large checkpoint for image captioning, served through Hugging Face Inference. Given a photo it returns a short English caption. The large variant gives more accurate captions than the base model and is a common drop-in for alt-text and image indexing.
blipcaptioningvision-understanding
Currently not offered
Meta
Meta Detectron2 object-detection and segmentation toolkit. Mask R-CNN, Cascade R-CNN, panoptic FPN and many other model variants in one wrapper.
segmentationvision-understandingdetection
Deactivated
Meta
Meta DINOv2 self-supervised vision backbone. Pretrained features for classification, segmentation and depth without task-specific fine-tuning.
segmentationvision-understandingopen-weights
Deactivated
Community
DWPose whole-body 2D pose estimator. Two-stage knowledge-distilled model with strong accuracy on face, hands and body keypoints simultaneously.
posevision-understandingopen-source
Deactivated
Google DeepMind
Google Gemini 1.5 Flash, the fast low-cost multimodal model. 1M-token context, image/audio/video input, good for high-volume captioning, classification and long-video skim tasks.
geminivisionmultimodal
Currently not offered
xAI's vision-capable Grok 2 snapshot. Image-in, text-out with strong multilingual instruction following.
vision
Deactivated
Community
Microsoft HRNet high-resolution pose-estimation backbone. Parallel multi-resolution streams yield strong accuracy on COCO keypoint benchmarks.
posevision-understandingmicrosoft
Deactivated
Community
OpenGVLab InternVL 2.5 78B. Open-source vision-language model approaching GPT-4o on MMMU, OCRBench and Math-Vista benchmarks.
multimodalvision-understandingopengvlab
Deactivated
Microsoft
Microsoft LayoutLMv3 multimodal document model. Unified text/image masking pretraining for form understanding, receipts and document QA.
ocrvision-understandingopen-source
Deactivated
Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.
llamamultimodalvision
Deactivated
Community
LLaVA v1.6 on a Nous-Hermes-2 34B base, served on Replicate. Open-source vision-language assistant for image question answering, description and visual reasoning at higher resolution.
llavavision-understandingopen-source
Deactivated
Community
LMMs-Lab LLaVA-OneVision 72B. Unified single-image, multi-image and video instruction-tuned VLM with task-transfer across modalities.
multimodalvision-understandingllava
Deactivated
Community
Marker PDF-to-Markdown conversion pipeline. Combines layout, OCR and equation models to produce clean Markdown with preserved tables and formulas.
ocrvision-understandingpdf
Currently not offered
Google DeepMind
Google MediaPipe Pose. Lightweight on-device-friendly 33-keypoint 3D pose estimator with optional segmentation mask output.
posevision-understandingopen-source
Deactivated
Mistral AI
Mistral OCR API. Document-understanding model with strong table and equation extraction, and structured JSON output.
ocrvision-understandingapi
Deactivated
Mistral AI
Mistral's 124B multimodal flagship. 123B decoder + 1B vision encoder, 128k ctx, up to 30 images per request.
pixtralmultimodalvision
Deactivated
Community
OpenMMLab MMPose toolbox. Wraps RTMPose, HRNet, HigherHRNet and many other pose models behind a unified inference API.
posevision-understandingmmpose
Deactivated
Microsoft
Microsoft Phi-3.5 Vision Instruct. Small (4.2B) multimodal model with strong document, OCR and multi-image reasoning at low cost.
multimodalvision-understandingopen-weights
Deactivated
Alibaba (Qwen)
Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.
multimodalvisionopen-weights
Deactivated
Alibaba (Qwen)
Alibaba Qwen2.5-VL 7B via Hugging Face Inference. Open-weights image-text-to-text model with improved OCR, chart and table reading, object grounding and long-document understanding.
vision-understandingopen-weightsocr
Currently not offered
Reka
Reka's frontier multimodal model supporting text, image, video and audio inputs.
multimodalvideo-understanding
Currently not offered
Reka
Reka's small on-device-friendly multimodal model. ~7B parameters, 16k context.
multimodaledgesmall
Currently not offered
Reka
Reka's 21B dense multimodal model balancing speed and quality. Up to 128k context.
multimodalcost-efficient
Currently not offered
Community
ETH Zurich SAM-HQ. High-quality mask refinement on top of SAM. Sharper edges and finer structure than the original Segment Anything model.
segmentationvision-understandingopen-source
Deactivated
Microsoft
Microsoft TrOCR large transformer-based OCR. End-to-end visual encoder plus text decoder, trained on synthetic and printed real-world data.
ocrvision-understandingopen-source
Deactivated
Community
ViTPose plain-vision-transformer pose estimator. State-of-the-art keypoint accuracy on MS-COCO with a minimal architecture.
posevision-understandingtransformer
Deactivated
01.AI
01.AI Yi-VL 34B vision-language model. Bilingual (CN/EN) image understanding, strong CMMMU and MMMU performance among open-weights VLMs.
multimodalvision-understanding01ai
Deactivated
Their pages stay online, but they canβt be run at the moment.
Multimodal models accept text plus images (sometimes plus audio or video) and produce text output. Reach for one when your input contains images and your output is structured information: extract an invoice, describe a chart, transcribe a handwritten note, answer questions about a UI screenshot.
Pricing is per-token like regular LLMs, with one twist: every image is counted as a fixed number of tokens β typically 250-1,500 tokens depending on resolution and detail mode. A standard 1024Γ1024 image costs roughly the same as a 1,000-word text input. High-detail mode (preserving fine text and small UI elements) costs 2-4Γ more. Plan budgets accordingly β a workload that processes thousands of receipt scans per day adds up quickly at flagship rates; multiply the per-token price on the model page by your image token count.
The trade-off is OCR accuracy, reasoning quality, and cost. Flagships (GPT-5 Vision, Claude 4.6, Gemini 2.5) read complex layouts and reason over chart contents very reliably. Specialized OCR-first models (Qwen 2.5 VL, Pixtral, InternVL) sometimes outperform on pure text extraction at a fraction of the cost. For pure document-to-JSON pipelines, a specialized model with a strict JSON schema usually wins. For document-and-reasoning workloads ('extract this invoice AND tell me if the tax math is right'), flagships win.
Watch out for resolution limits: most models downscale very large images before processing, which can destroy fine text in screenshots and dense documents. Pre-process β slice tall documents into single-page images, upscale low-resolution scans before sending β to preserve readability. Also watch out for hallucinations on hard-to-read regions; multimodal models tend to confidently invent text where the source is illegible.
Top picks above cover the most accurate flagship, the cheapest workhorse, the highest-resolution supporter, and the fastest streaming option.
GPT-5 Vision and Claude 4.6 lead on complex layouts, charts, and handwriting. For pure OCR on clean documents, Qwen 2.5 VL and Pixtral often match flagship quality at one-tenth the cost. Run a sample of your own documents on each before committing.
Every image is converted to a fixed token count β typically 250-1,500 tokens depending on resolution and detail mode. So a 1024Γ1024 image costs the same as ~1,000 tokens of text. Multi-page documents multiply linearly; check each model card for exact image-token rates.
JPEG, PNG, WebP, and GIF (first frame) are universal. Some models also accept PDF directly (one image per page). Max file size is typically 20 MB; max dimensions 8K on a side. Pre-resize larger files to avoid silent downscaling.
Flagships handle clear handwriting reliably and most cursive scripts in major languages. Quality drops on heavily stylized writing, faded ink, and complex shorthand. For high-volume handwriting workloads, fine-tune on a sample of your domain or pair with a specialized OCR pre-processor.
Yes β chart reading is one of the strongest use-cases. Flagships extract data points from bar charts, line graphs, and pie charts; they reason about axis scales and identify trends. For dense scientific figures, ask for structured output (CSV or JSON) rather than narrative description.
Quality degrades sharply below 150 DPI or above 30% noise. Pre-process with deskewing, contrast normalization, and lightweight denoising. For very poor scans, run an upscaler first (from /models/image) to restore readability before sending.
Yes β most flagships accept 10-20 images in a single prompt. Useful for comparing screenshots, processing multi-page documents in one pass, or analyzing sequences of frames. Total token count (including all image tokens) must fit inside the context window.
Some multimodal models (Gemini 2.5 Pro, Qwen 2.5 VL) accept video directly and sample frames internally. For others, extract frames yourself (every 1-2 seconds is typical) and send them as a multi-image prompt. For real-time video analysis, look at dedicated video-understanding models.
Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.