Qwen2-VL-72B Instruct

MultimodalNicht verfügbar
von Alibaba / QwenModell-ID: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Status
Nicht verfügbar
Kontext
32,768 Token
Max. Ausgabe
8,192 Token
Eingabe → Ausgabe
Text + Bild + Video → Text
Entwickler
Alibaba / Qwen
Aktualisiert
25. Juni 2026

Qwen2-VL-72B Instruct ist derzeit nicht verfügbar

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert.

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

Zu den Alternativen
01

Vergleichbare Modelle

Alle dieser Kategorie
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ca. $ 0.00030/Lauf

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    $ 6.00/1 Mio. In

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    $ 3.60/1 Mio. In

02

Playground

Qwen2-VL-72B Instruct ausprobieren

Chat

Derzeit nicht verfügbar

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert.

Der Playground ist deaktiviert. Vergleichbare Modelle findest du in derselben Kategorie: Alternativen ansehen

Qwen2-VL-72B Instruct ausprobieren

Schick eine Nachricht. Die Antwort kommt vollständig, sobald das Modell fertig ist (ohne Streaming).

System-Prompt
Max. Antwortlänge (Token)

Dieser Lauf

Kein Preis – derzeit nicht verfügbar.

Neu hier?

5 Gratis-Credits ($ 0.05) bei Anmeldung mit Google

Nutzbar 24 Stunden nach der Anmeldung, bis zu 5 Läufe pro Tag und höchstens 2 Credits je Lauf. Andere Anmeldearten starten ohne Guthaben.

03

Über Qwen2-VL-72B Instruct

Kurz gesagtStand: 25. Juni 2026

Qwen2-VL-72B Instruct ist ein Modell von Alibaba / Qwen aus der Kategorie Multimodal. Über Railwail ist Qwen2-VL-72B Instruct derzeit nicht verfügbar. Das Kontextfenster umfasst 32,768 Token, eine Antwort bis zu 8,192 Token.

Hintergrund

Über Alibaba DAMO Academy (Qwen Team)

Gegründet 2017 · Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Alibaba DAMO Academy (Qwen Team) besuchen

Architektur

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Parameter
72B (~73B with vision encoder)
Kontext
131,072 Token

Fähigkeiten

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Training & Lizenz

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

Lizenz: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Sicherheitstests: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Bekannte Grenzen

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Preise

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

05

API

Rufe Qwen2-VL-72B Instruct mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Derzeit nicht verfügbar

Das Modell hat keinen geprüften Preis oder ist deaktiviert; API-Aufrufe werden abgelehnt.

06

Spezifikationen

Modell-ID
qwen2-vl-72b-instruct
Entwickler
Alibaba / Qwen
Kategorie
Multimodal
Eingabe
Text, Bild, Video
Ausgabe
Text
Kontextfenster
32,768 Token
Max. Ausgabe
8,192 Token
Modellgrösse
72B (~73B with vision encoder)
Lizenz
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Katalogeintrag aktualisiert
25. Juni 2026

Schlagwörter

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Einsatzgebiete

Wofür es genutzt wird

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Häufige Fragen

Was ist Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct ist ein Modell von Alibaba / Qwen aus der Kategorie Multimodal. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet Qwen2-VL-72B Instruct bei Railwail?

Qwen2-VL-72B Instruct lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie gross ist das Kontextfenster von Qwen2-VL-72B Instruct?

Das Kontextfenster von Qwen2-VL-72B Instruct umfasst 32,768 Token. Eine Antwort kann bis zu 8,192 Token lang sein.

Wie schnell ist Qwen2-VL-72B Instruct?

Für Qwen2-VL-72B Instruct gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist Qwen2-VL-72B Instruct besser als BLIP?

Das hängt von der Aufgabe ab. Qwen2-VL-72B Instruct (Alibaba / Qwen) und BLIP (Salesforce) sind beide Modelle aus der Kategorie Multimodal. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

Qwen2-VL-72B Instruct und BLIP vergleichen

Kann Qwen2-VL-72B Instruct Bilder verarbeiten?

Ja. Qwen2-VL-72B Instruct nimmt neben Text auch Bilder als Eingabe an.

Kann ich Qwen2-VL-72B Instruct gerade nutzen?

Derzeit nicht verfügbar: Dieses Modell ist deaktiviert. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0.01.