Qwen2-VL-72B Instruct

MultimodalIndisponibil
de Alibaba / QwenID model: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Status
Indisponibil
Context
32.768 tokeni
Ieșire max.
8.192 tokeni
Intrare → ieșire
Text + Imagine + Video → Text
Dezvoltator
Alibaba / Qwen
Actualizat
25 iunie 2026

Qwen2-VL-72B Instruct nu este disponibil în acest moment

Indisponibil în prezent: acest model a fost dezactivat.

Poți citi în continuare detaliile pe această pagină. Alege una dintre alternativele disponibile de mai jos pentru a rula imediat un model comparabil.

Mergi la alternative
01
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/rulare

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M in

02

Playground

Încearcă Qwen2-VL-72B Instruct

Chat

Indisponibil în prezent

Indisponibil în prezent: acest model a fost dezactivat.

Playground-ul este dezactivat. Modele comparabile găsești în aceeași categorie: Vezi alternativele

Încearcă Qwen2-VL-72B Instruct

Trimite un mesaj. Răspunsul sosește complet când modelul termină (fără streaming).

Prompt de sistem
Lungimea max. a răspunsului (tokeni)

Această rulare

Fără preț – momentan indisponibil.

Nou aici?

5 credite gratuite (0,05 USD) când te înregistrezi cu Google

Utilizabil 24 ore după înregistrare, până la 5 rulări pe zi și maximum 2 credite pe rulare. Alte metode de conectare încep fără credite.

03

Despre Qwen2-VL-72B Instruct

Pe scurtDin 25 iunie 2026

Qwen2-VL-72B Instruct este un model de Alibaba / Qwen din categoria Multimodal. Qwen2-VL-72B Instruct nu este disponibil în prezent pe Railwail. Fereastra de context conține 32.768 token-uri, iar un răspuns poate fi lung de până la 8.192 token-uri.

Fundal

Despre Alibaba DAMO Academy (Qwen Team)

Fondat 2017 · Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Vizitează Alibaba DAMO Academy (Qwen Team)

Arhitectură

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Parametri
72B (~73B with vision encoder)
Context
131.072 tokeni

Capabilități

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Antrenament & licență

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

Licență: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Teste de siguranță: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Limitări cunoscute

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Prețuri

Indisponibil în prezent: acest model a fost dezactivat. Nu există preț pentru acest model în acest moment, deci nu poate fi executat.

05

API

Apelează Qwen2-VL-72B Instruct cu cheia ta API Railwail. Folosește acest ID de model în cerere:

Indisponibil în prezent

Modelul nu are un preț verificat sau este dezactivat; apelurile API sunt refuzate.

06

Specificații

ID model
qwen2-vl-72b-instruct
Dezvoltator
Alibaba / Qwen
Categorie
Multimodal
Intrare
Text, Imagine, Video
Ieșire
Text
Fereastră de context
32.768 tokeni
Ieșire max.
8.192 tokeni
Dimensiune model
72B (~73B with vision encoder)
Licență
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Intrare catalog actualizată
25 iunie 2026

Etichete

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Cazuri de utilizare

Pentru ce se folosește

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Întrebări frecvente

Ce este Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct este un model de Alibaba / Qwen din categoria Multimodal. Este listat pe Railwail, dar nu poate fi rulat în acest moment.

Cât costă Qwen2-VL-72B Instruct pe Railwail?

Qwen2-VL-72B Instruct nu poate fi rulat pe Railwail în acest moment, deci nu există preț curent. Alternativele disponibile cu prețuri sunt listate mai jos pe această pagină.

Care este fereastra de context a Qwen2-VL-72B Instruct?

Fereastra de context a Qwen2-VL-72B Instruct conține 32.768 token-uri. Un răspuns poate fi lung de până la 8.192 token-uri.

Cât de rapid este Qwen2-VL-72B Instruct?

Nu sunt suficiente rulări măsurate ale Qwen2-VL-72B Instruct pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.

Este Qwen2-VL-72B Instruct mai bun decât BLIP?

Depinde de sarcină. Qwen2-VL-72B Instruct (Alibaba / Qwen) și BLIP (Salesforce) sunt ambele modele din categoria Multimodal. Pagina de comparație arată prețurile și specificațiile lor una lângă alta.

Compară Qwen2-VL-72B Instruct și BLIP

Poate Qwen2-VL-72B Instruct procesa imagini?

Da. Qwen2-VL-72B Instruct acceptă imagini ca intrare, pe lângă text.

Pot folosi Qwen2-VL-72B Instruct chiar acum?

Indisponibil în prezent: acest model a fost dezactivat. Pagina rămâne online; alternativele disponibile din aceeași categorie sunt listate mai jos.

Toate modelele printr-o singură API

O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.