Qwen2-VL-72B Instruct

МультимодальніНедоступно
від Alibaba / QwenID моделі: qwen2-vl-72b-instruct

Alibaba's 72B vision-language model with M-RoPE and dynamic resolution. Strong document and video understanding.

Статус
Недоступно
Контекст
32,768 токенів
Макс. вихід
8,192 токенів
Вхід → вихід
Текст + Зображення + Відео → Текст
Розробник
Alibaba / Qwen
Оновлено
25 червня 2026 р.

Qwen2-VL-72B Instruct наразі недоступна

Наразі недоступно: цю модель деактивовано.

Ви можете прочитати деталі на цій сторінці. Виберіть одну з доступних альтернатив нижче, щоб одразу запустити порівняну модель.

До альтернатив
01

Порівнювані моделі

Усі в цій категорії
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/запуск

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M вх

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M вх

02

Playground

Спробувати Qwen2-VL-72B Instruct

Чат

Наразі недоступна

Наразі недоступно: цю модель деактивовано.

Playground вимкнено. Порівняльні моделі ви знайдете в тій же категорії: Переглянути альтернативи

Спробувати Qwen2-VL-72B Instruct

Надішліть повідомлення. Відповідь надійде повністю, коли модель завершить роботу (без потокової передачі).

Системний промпт
Макс. довжина відповіді (токени)

Цей запуск

Без ціни – наразі недоступно.

Новенький?

5 безплатних кредитів (0,05 USD) при реєстрації через Google

Доступно 24 годин після реєстрації, до 5 запусків на день і максимум 2 кредитів за запуск. Інші способи входу стартують без кредитів.

03

Про Qwen2-VL-72B Instruct

КороткоСтаном на 25 червня 2026 р.

Qwen2-VL-72B Instruct — це модель від Alibaba / Qwen у категорії Мультимодальні. На Railwail Qwen2-VL-72B Instruct наразі недоступна. Контекстне вікно містить 32,768 токенів, а одна відповідь може бути довжиною до 8,192 токенів.

Фон

Про Alibaba DAMO Academy (Qwen Team)

Засновано 2017 · Hangzhou, China

The Qwen (Tongyi Qianwen) team sits inside Alibaba Cloud's DAMO Academy, the company's research arm founded in 2017 in Hangzhou. The team is led by Junyang Lin and Le Hou and counts dozens of researchers across NLP, vision and speech. Qwen has produced one of the most prolific open-source model lines in the world, including Qwen-1.5, Qwen2 (June 2024), Qwen2.5 (September 2024), the Code, Math, Audio and VL (vision-language) families, and the December 2024 release of Qwen2.5-VL. Qwen2-VL launched in August 2024 in 2B, 7B and 72B sizes, all released on Hugging Face and ModelScope; the 72B Instruct variant became one of the top open-weights vision-language models worldwide, frequently matching closed-source peers on OCR-heavy benchmarks like DocVQA and ChartQA. Alibaba offers Qwen models commercially through Alibaba Cloud and Bailian.

Відвідати Alibaba DAMO Academy (Qwen Team)

Архітектура

Decoder-only Transformer with Naive Dynamic Resolution Vision Transformer

Qwen2-VL-72B-Instruct combines the Qwen2 72B decoder-only Transformer with a custom 675M ViT vision encoder using Naive Dynamic Resolution: instead of resizing every image to a fixed grid, the encoder accepts the native resolution and generates a variable number of visual tokens per image. The model also introduces Multimodal Rotary Position Embedding (M-RoPE) that encodes positions in time (for video), height and width separately, enabling single-stream multimodal video understanding. The model supports up to 20 minutes of video input via uniform frame sampling, single-frame image input at variable resolution up to ~16K visual tokens, and a 131,072-token text context window. Training proceeded in three stages: contrastive vision-language pretraining, multimodal pretraining on interleaved image-text and video-text data, and supervised fine-tuning with chain-of-thought multimodal instructions. Weights are released under the Qwen licence (free for commercial use under specific terms).

Параметри
72B (~73B with vision encoder)
Контекст
131,072 токенів

Можливості

  • Open-weights 72B vision-language model under permissive Qwen licence
  • Naive Dynamic Resolution: native image aspect ratio without fixed grid
  • Multimodal Rotary Position Embedding (M-RoPE) for joint image and video
  • Up to 20 minutes of video understanding
  • 131K-token text context
  • Top open-weights scores on DocVQA, ChartQA, MathVista, RealWorldQA
  • Strong OCR across English, Chinese, Japanese, Korean and European languages
  • Best for: open-weights document AI, video QA, OCR-heavy multilingual workloads

Навчання та ліцензія

Multi-stage curriculum: contrastive vision-language pretraining on large web image-text pairs, multimodal pretraining on interleaved image-text and video-text data, supervised fine-tuning on curated chain-of-thought multimodal instructions.

Ліцензія: Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).

Тестування безпеки: Alibaba publishes a model card with safety evaluations and integrates Tongyi safety filters in cloud deployments; no separate full red-team report.

Відомі обмеження

  • Serving 72B requires multi-GPU infrastructure
  • Video understanding limited to 20 minutes uniform sampling
  • Hallucination on extreme OCR cases
  • Licence has MAU and competing-services restrictions
  • Audio input requires separate Qwen-Audio model
04

Ціни

Наразі недоступно: цю модель деактивовано. Наразі для цієї моделі немає ціни, тому її не можна запустити.

05

API

Викличте Qwen2-VL-72B Instruct за допомогою вашого API-ключа Railwail. Використовуйте цей ID моделі в запиті:

Наразі недоступно

Модель не має перевіреної ціни або деактивована; виклики API відхиляються.

06

Характеристики

ID моделі
qwen2-vl-72b-instruct
Розробник
Alibaba / Qwen
Вхідні дані
Текст, Зображення, Відео
Вихідні дані
Текст
Контекстне вікно
32,768 токенів
Макс. вихідні дані
8,192 токенів
Розмір моделі
72B (~73B with vision encoder)
Ліцензія
Qwen Licence (commercial use permitted under 100M MAU; bespoke licence required above).
Запис у каталозі оновлено
25 червня 2026 р.

Теги

  • qwen
  • alibaba
  • multimodal
  • vision
  • open-weights
  • video-understanding
  • pricing-tbd
07

Варіанти використання

Для чого це використовується

  • Open-weights document AI for multilingual OCR
  • Video question answering up to 20 minutes
  • Chart and diagram understanding for analytics
  • Chinese / Japanese / Korean OCR-heavy workloads
  • Multimodal AI assistants in Chinese cloud regions
08

Часті питання

Що таке Qwen2-VL-72B Instruct?

Qwen2-VL-72B Instruct — це модель від Alibaba / Qwen у категорії Мультимодальні. Вона внесена до каталогу Railwail, але наразі не може бути запущена.

Скільки коштує Qwen2-VL-72B Instruct на Railwail?

Qwen2-VL-72B Instruct наразі не можна запустити на Railwail, тому поточної ціни немає. Доступні альтернативи з цінами наведені далі на цій сторінці.

Яке контекстне вікно у Qwen2-VL-72B Instruct?

Контекстне вікно Qwen2-VL-72B Instruct містить 32,768 токенів. Одна відповідь може бути довжиною до 8,192 токенів.

Наскільки швидкий Qwen2-VL-72B Instruct?

На Railwail поки що недостатньо виміряних запусків Qwen2-VL-72B Instruct, щоб вказати час виконання. Це залежить від введення, налаштувань та навантаження у постачальника.

Чи Qwen2-VL-72B Instruct краще за BLIP?

Це залежить від завдання. Qwen2-VL-72B Instruct (Alibaba / Qwen) і BLIP (Salesforce) — обидві моделі в категорії Мультимодальні. На сторінці порівняння показані їхні ціни та специфікації поруч.

Порівняти Qwen2-VL-72B Instruct і BLIP

Чи може Qwen2-VL-72B Instruct обробляти зображення?

Так. Qwen2-VL-72B Instruct приймає зображення як введення, крім тексту.

Чи можу я використовувати Qwen2-VL-72B Instruct прямо зараз?

Наразі недоступно: цю модель деактивовано. Сторінка залишається в мережі; доступні альтернативи з тієї ж категорії наведені далі.

Усі моделі через один API

Один API-ключ для всіх моделей на Railwail. Використання оплачується з попередньо поповнених кредитів, 1 кредит = 0,01 USD.