Qwen2-VL-72B Instruct

MultimodalUnavailable
by Alibaba / QwenModel ID: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Status
Unavailable
Context
32,768 tokens
Max. output
8,192 tokens
Input β†’ output
Text + Image + Video β†’ Text
Developer
Alibaba / Qwen
Updated
June 25, 2026

Qwen2-VL-72B Instruct is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    β‰ˆ $0.00030/run

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    $6.00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    $3.60/1M in

02

Playground

Try Qwen2-VL-72B Instruct

Chat

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Qwen2-VL-72B Instruct

Send a message. The answer arrives in full once the model is done (no streaming).

System prompt
Max. answer length (tokens)

This run

No price – currently unavailable.

New here?

5 free credits ($0.05) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Qwen2-VL-72B Instruct

TL;DRAs of June 25, 2026

Qwen2-VL-72B Instruct is a model by Alibaba / Qwen in the Multimodal category. Qwen2-VL-72B Instruct is currently not available on Railwail. The context window holds 32,768 tokens, and one response can be up to 8,192 tokens long.

Background

About Alibaba DAMO Academy (Qwen Team)

Founded 2017 Β· Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Visit Alibaba DAMO Academy (Qwen Team)

Architecture

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Parameters
72B (~73B with vision encoder)
Context
131,072 tokens

Capabilities

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Training & license

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

License: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Safety testing: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Known limitations

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call Qwen2-VL-72B Instruct with your Railwail API key. Use this model ID in the request:
qwen2-vl-72b-instructAPI documentationGet an API key

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
qwen2-vl-72b-instruct
Developer
Alibaba / Qwen
Category
Multimodal
Input
Text, Image, Video
Output
Text
Context window
32,768 tokens
Max. output
8,192 tokens
Model size
72B (~73B with vision encoder)
License
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Catalog entry updated
June 25, 2026

Tags

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Use cases

What it is used for

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Frequently asked questions

What is Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct is a model by Alibaba / Qwen in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does Qwen2-VL-72B Instruct cost on Railwail?

Qwen2-VL-72B Instruct cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Qwen2-VL-72B Instruct?

The context window of Qwen2-VL-72B Instruct holds 32,768 tokens. One response can be up to 8,192 tokens long.

How fast is Qwen2-VL-72B Instruct?

There are not enough measured runs of Qwen2-VL-72B Instruct on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Qwen2-VL-72B Instruct better than BLIP?

That depends on the task. Qwen2-VL-72B Instruct (Alibaba / Qwen) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Qwen2-VL-72B Instruct and BLIP

Can Qwen2-VL-72B Instruct process images?

Yes. Qwen2-VL-72B Instruct accepts images as input in addition to text.

Can I use Qwen2-VL-72B Instruct right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.