Qwen2-VL-72B Instruct

MultimodalneNiedostępne
od Alibaba / QwenID modelu: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Status
Niedostępne
Kontekst
32 768 tokenów
Maks. wyjście
8192 tokenów
Wejście → Wyjście
Tekst + Obraz + Wideo → Tekst
Deweloper
Alibaba / Qwen
Zaktualizowano
25 czerwca 2026

Qwen2-VL-72B Instruct jest obecnie niedostępny

Obecnie niedostępne: ten model został dezaktywowany.

Możesz nadal przeczytać szczegóły na tej stronie. Wybierz jedną z dostępnych alternatyw poniżej, aby od razu uruchomić porównywalny model.

Przejdź do alternatyw
01

Porównywalne modele

Wszystkie w tej kategorii
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/uruchomienie

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M wej.

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M wej.

02

Playground

Spróbuj Qwen2-VL-72B Instruct

Chat

Niedostępny

Obecnie niedostępne: ten model został dezaktywowany.

Plac zabaw jest wyłączony. Porównywalne modele znajdziesz w tej samej kategorii: Przeglądaj alternatywy

Spróbuj Qwen2-VL-72B Instruct

Wyślij wiadomość. Odpowiedź pojawi się w całości, gdy model będzie gotowy (bez streamingu).

Prompt systemowy
Maks. długość odpowiedzi (tokeny)

To uruchomienie

Brak ceny – obecnie niedostępne.

Nowy tutaj?

5 darmowych kredytów (0,05 USD) po zarejestrowaniu się przez Google

Dostępne 24 godzin po rejestracji, do 5 uruchomień dziennie i maksymalnie 2 kredytów na uruchomienie. Inne metody logowania uruchamiają się bez kredytów.

03

O Qwen2-VL-72B Instruct

Krótko mówiącStan na 25 czerwca 2026

Qwen2-VL-72B Instruct to model opracowany przez Alibaba / Qwen w kategorii Multimodalne. Qwen2-VL-72B Instruct nie jest obecnie dostępny w serwisie Railwail. Okno kontekstu zawiera 32 768 tokenów, a jedna odpowiedź może mieć do 8192 tokenów.

Tło

O Alibaba DAMO Academy (Qwen Team)

Założona 2017 · Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Odwiedź Alibaba DAMO Academy (Qwen Team)

Architektura

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Parametry
72B (~73B with vision encoder)
Kontekst
131 072 tokenów

Możliwości

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Trening i licencja

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

Licencja: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Testy bezpieczeństwa: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Znane ograniczenia

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Ceny

Obecnie niedostępne: ten model został dezaktywowany. Dla tego modelu nie ma ceny w tej chwili, dlatego nie można go uruchomić.

05

API

Wywołaj Qwen2-VL-72B Instruct za pomocą klucza API Railwail. Użyj tego ID modelu w żądaniu:

Obecnie niedostępne

Model nie ma zweryfikowanej ceny lub jest wyłączony; wywołania API są odrzucane.

06

Specyfikacje

ID modelu
qwen2-vl-72b-instruct
Deweloper
Alibaba / Qwen
Kategoria
Multimodalne
Wejście
Tekst, Obraz, Wideo
Wyjście
Tekst
Okno kontekstu
32 768 tokenów
Maks. wyjście
8192 tokenów
Rozmiar modelu
72B (~73B with vision encoder)
Licencja
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Wpis w katalogu zaktualizowany
25 czerwca 2026

Tagi

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Przypadki użycia

Do czego się go używa

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Często zadawane pytania

Co to jest Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct to model opracowany przez Alibaba / Qwen w kategorii Multimodalne. Jest wymieniony w katalogu Railwail, ale nie może być uruchomiony w tej chwili.

Ile kosztuje Qwen2-VL-72B Instruct w serwisie Railwail?

Qwen2-VL-72B Instruct nie może być uruchomiony w serwisie Railwail w tej chwili, dlatego nie ma aktualnej ceny. Dostępne alternatywy z cenami są wymienione poniżej na tej stronie.

Jakie jest okno kontekstu Qwen2-VL-72B Instruct?

Okno kontekstu Qwen2-VL-72B Instruct zawiera 32 768 tokenów. Jedna odpowiedź może mieć do 8192 tokenów.

Jak szybki jest Qwen2-VL-72B Instruct?

Dla Qwen2-VL-72B Instruct jest jeszcze zbyt mało zmierzonych przebiegów w serwisie Railwail, aby podać czas przebiegu. Zależy to od wejścia, ustawień i obciążenia u dostawcy.

Czy Qwen2-VL-72B Instruct jest lepszy niż BLIP?

To zależy od zadania. Qwen2-VL-72B Instruct (Alibaba / Qwen) i BLIP (Salesforce) to oba modele z kategorii Multimodalne. Strona porównania pokazuje ich ceny i specyfikacje obok siebie.

Porównaj Qwen2-VL-72B Instruct i BLIP

Czy Qwen2-VL-72B Instruct może przetwarzać obrazy?

Tak. Qwen2-VL-72B Instruct akceptuje obrazy jako wejście oprócz tekstu.

Czy mogę używać Qwen2-VL-72B Instruct teraz?

Obecnie niedostępne: ten model został dezaktywowany. Strona pozostaje online; dostępne alternatywy z tej samej kategorii są wymienione poniżej.

Wszystkie modele przez jedno API

Jeden klucz API dla każdego modelu na Railwail. Opłaty pobierane są z przedpłaconych kredytów, 1 kredyt = 0,01 USD.