Qwen2-VL-72B Instruct

MultimodálneNedostupné
od Alibaba / QwenID modelu: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Stav
Nedostupné
Kontext
32 768 tokenov
Max. výstup
8 192 tokenov
Vstup → výstup
Text + Obrázok + Video → Text
Vývojár
Alibaba / Qwen
Aktualizované
25. júna 2026

Qwen2-VL-72B Instruct nie je momentálne dostupný

Momentálne nedostupné: tento model bol deaktivovaný.

Podrobnosti na tejto stránke si môžete prečítať. Vyberte si jednu z dostupných alternatív nižšie a spustite porovnateľný model hneď.

Prejsť na alternatívy
01

Porovnateľné modely

Všetky v tejto kategórii
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/spustenie

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M vstup

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M vstup

02

Playground

Vyskúšajte Qwen2-VL-72B Instruct

Chat

Momentálne nedostupné

Momentálne nedostupné: tento model bol deaktivovaný.

Playground je vypnutý. Porovnateľné modely nájdete v tej istej kategórii: Pozrieť alternatívy

Vyskúšajte Qwen2-VL-72B Instruct

Pošli správu. Odpoveď príde úplne, keď je model hotový (bez streamovania).

Systémový prompt
Maximálna dĺžka odpovede (tokeny)

Tento beh

Bez ceny – momentálne nedostupné.

Nový tu?

5 bezplatných credits (0,05 USD) pri registrácii cez Google

Použiteľné 24 hodín po registrácii, až 5 behov za deň a maximálne 2 credits za beh. Ostatné spôsoby prihlásenia sa spúšťajú bez credits.

03

O Qwen2-VL-72B Instruct

StručneK 25. júna 2026

Qwen2-VL-72B Instruct je model od Alibaba / Qwen v kategórii Multimodálne. Qwen2-VL-72B Instruct nie je v súčasnosti dostupný na Railwail. Kontextné okno obsahuje 32 768 tokenov a jedna odpoveď môže byť dlhá až 8 192 tokenov.

Pozadie

O Alibaba DAMO Academy (Qwen Team)

Založené 2017 · Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Navštíviť Alibaba DAMO Academy (Qwen Team)

Architektúra

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Parametre
72B (~73B with vision encoder)
Kontext
131 072 tokenov

Schopnosti

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Tréning a licencia

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

Licencia: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Bezpečnostné testy: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Známe obmedzenia

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Ceny

Momentálne nedostupné: tento model bol deaktivovaný. V súčasnosti nie je cena za tento model, preto ho nie je možné spustiť.

05

API

Zavolajte Qwen2-VL-72B Instruct s vaším API kľúčom Railwail. V požiadavke použite toto ID modelu:

Momentálne nedostupné

Model nemá overenú cenu alebo je deaktivovaný; volania API sú odmietnuté.

06

Špecifikácie

ID modelu
qwen2-vl-72b-instruct
Vývojár
Alibaba / Qwen
Kategória
Multimodálne
Vstup
Text, Obrázok, Video
Výstup
Text
Kontextové okno
32 768 tokenov
Max. výstup
8 192 tokenov
Veľkosť modelu
72B (~73B with vision encoder)
Licencia
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Katalógová položka aktualizovaná
25. júna 2026

Značky

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Prípady použitia

Na čo sa používa

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Často kladené otázky

Čo je Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct je model od Alibaba / Qwen v kategórii Multimodálne. Je uvedený na Railwail, ale momentálne ho nie je možné spustiť.

Koľko stojí Qwen2-VL-72B Instruct na Railwail?

Qwen2-VL-72B Instruct nie je momentálne možné spustiť na Railwail, preto nie je aktuálna cena. Dostupné alternatívy s cenami sú uvedené nižšie na tejto stránke.

Aké je kontextné okno Qwen2-VL-72B Instruct?

Kontextné okno Qwen2-VL-72B Instruct obsahuje 32 768 tokenov. Jedna odpoveď môže byť dlhá až 8 192 tokenov.

Ako rýchly je Qwen2-VL-72B Instruct?

Pre Qwen2-VL-72B Instruct je na Railwail zatiaľ príliš málo meraných spustení na určenie doby spustenia. Závisí to od vstupu, nastavení a zaťaženia u poskytovateľa.

Je Qwen2-VL-72B Instruct lepší ako BLIP?

Závisí to od úlohy. Qwen2-VL-72B Instruct (Alibaba / Qwen) a BLIP (Salesforce) sú oba modely v kategórii Multimodálne. Stránka porovnania zobrazuje ich ceny a špecifikácie vedľa seba.

Porovnať Qwen2-VL-72B Instruct a BLIP

Môže Qwen2-VL-72B Instruct spracovávať obrázky?

Áno. Qwen2-VL-72B Instruct akceptuje obrázky ako vstup okrem textu.

Môžem Qwen2-VL-72B Instruct používať práve teraz?

Momentálne nedostupné: tento model bol deaktivovaný. Stránka zostáva online; dostupné alternatívy z tej istej kategórie sú uvedené nižšie.

Všetky modely cez jedno API

Jeden API kľúč pre všetky modely na Railwail. Použitie sa účtuje z predplateného kreditu, 1 kredit = 0,01 USD.