Qwen2-VL-72B Instruct

MultimodaleNon disponibile
di Alibaba / QwenID modello: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Stato
Non disponibile
Contesto
32.768 token
Max. output
8192 token
Input → output
Testo + Immagine + Video → Testo
Sviluppatore
Alibaba / Qwen
Aggiornato
25 giugno 2026

Qwen2-VL-72B Instruct non è attualmente disponibile

Attualmente non disponibile: questo modello è stato disattivato.

Puoi comunque leggere i dettagli su questa pagina. Scegli una delle alternative disponibili di seguito per eseguire subito un modello comparabile.

Vai alle alternative
01

Modelli comparabili

Tutti in questa categoria
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/esecuzione

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M in

02

Playground

Prova Qwen2-VL-72B Instruct

Chat

Attualmente non disponibile

Attualmente non disponibile: questo modello è stato disattivato.

Il playground è disabilitato. Trovi modelli comparabili nella stessa categoria: Visualizza alternative

Prova Qwen2-VL-72B Instruct

Invia un messaggio. La risposta arriva completa quando il modello ha finito (senza streaming).

Prompt di sistema
Lunghezza massima della risposta (token)

Questa esecuzione

Nessun prezzo – attualmente non disponibile.

Nuovo qui?

5 crediti gratuiti (0,05 USD) quando ti iscrivi con Google

Utilizzabile 24 ore dopo l'iscrizione, fino a 5 esecuzioni al giorno e al massimo 2 crediti per esecuzione. Altri metodi di accesso iniziano senza crediti.

03

Informazioni su Qwen2-VL-72B Instruct

RiassuntoA partire da 25 giugno 2026

Qwen2-VL-72B Instruct è un modello di Alibaba / Qwen nella categoria Multimodale. Qwen2-VL-72B Instruct non è attualmente disponibile su Railwail. La finestra di contesto contiene 32.768 token e una risposta può essere lunga fino a 8192 token.

Sfondo

Informazioni su Alibaba DAMO Academy (Qwen Team)

Fondato 2017 · Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Visita Alibaba DAMO Academy (Qwen Team)

Architettura

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Parametri
72B (~73B with vision encoder)
Contesto
131.072 token

Capacità

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Addestramento e licenza

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

Licenza: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Test di sicurezza: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Limitazioni note

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Prezzi

Attualmente non disponibile: questo modello è stato disattivato. Al momento non c'è un prezzo per questo modello, quindi non può essere eseguito.

05

API

Chiama Qwen2-VL-72B Instruct con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Attualmente non disponibile

Il modello non ha un prezzo verificato o è disattivato; le chiamate API vengono rifiutate.

06

Specifiche

ID modello
qwen2-vl-72b-instruct
Sviluppatore
Alibaba / Qwen
Categoria
Multimodale
Input
Testo, Immagine, Video
Output
Testo
Finestra di contesto
32.768 token
Output massimo
8192 token
Dimensione del modello
72B (~73B with vision encoder)
Licenza
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Voce di catalogo aggiornata
25 giugno 2026

Etichette

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Casi d'uso

A cosa serve

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Domande frequenti

Cos'è Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct è un modello di Alibaba / Qwen nella categoria Multimodale. È elencato su Railwail ma non può essere eseguito al momento.

Quanto costa Qwen2-VL-72B Instruct su Railwail?

Qwen2-VL-72B Instruct non può essere eseguito su Railwail al momento, quindi non c'è un prezzo attuale. Le alternative disponibili con i prezzi sono elencate più in basso in questa pagina.

Qual è la finestra di contesto di Qwen2-VL-72B Instruct?

La finestra di contesto di Qwen2-VL-72B Instruct contiene 32.768 token. Una risposta può essere lunga fino a 8192 token.

Quanto è veloce Qwen2-VL-72B Instruct?

Non ci sono ancora abbastanza esecuzioni misurate di Qwen2-VL-72B Instruct su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

Qwen2-VL-72B Instruct è migliore di BLIP?

Dipende dall'attività. Qwen2-VL-72B Instruct (Alibaba / Qwen) e BLIP (Salesforce) sono entrambi modelli nella categoria Multimodale. La pagina di confronto mostra i loro prezzi e le specifiche affiancati.

Confronta Qwen2-VL-72B Instruct e BLIP

Qwen2-VL-72B Instruct può elaborare immagini?

Sì. Qwen2-VL-72B Instruct accetta immagini come input oltre al testo.

Posso usare Qwen2-VL-72B Instruct adesso?

Attualmente non disponibile: questo modello è stato disattivato. La pagina rimane online; le alternative disponibili della stessa categoria sono elencate più in basso.

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.