Qwen2-VL 7B Instruct

MultimodalDisponible
de CommunityID del modelo: qwen2-vl-7b-instruct-rep

Alibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.

Precio
≈ 0,0024 US$/ejecución
Entrada → Salida
Texto + Imagen + Vídeo → Texto
Desarrollador
Community
Actualizado
23 de septiembre de 2026
01

Playground

Probar Qwen2-VL 7B Instruct

Entrada y salida

≈ 0,0024 US$/ejecución
Probar Qwen2-VL 7B Instruct

0 / 16.000

Configuración avanzada (2)
Resultado
La respuesta aparece aquí.

Esta ejecución

aprox. 0,0024 US$ · 0,24 créditos

Se reservan 0,0072 US$ (0,72 créditos) al inicio; se factura el tiempo real de GPU.

¿Eres nuevo aquí?

10 créditos gratis (0,10 US$) cuando te registres con Google

Disponible 24 horas después del registro, hasta 5 ejecuciones por día y como máximo 2 créditos por ejecución. Otros métodos de inicio de sesión comienzan sin créditos. Suficiente para 13 ejecuciones de este modelo.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    Describe this video in detail.

    Response

    The video features a monkey riding a skateboard in a park. The monkey is wearing roller skates and is seen moving forward on the skateboard, with its arms outstretched for balance. The park is filled with people sitting on benches and walking around, enjoying the sunny day. The monkey appears to be enjoying itself, and the scene is filled with a sense of fun and playfulness. The park is surrounded by trees and greenery, creating a peaceful and relaxing atmosphere. Overall, the video captures a unique and entertaining moment of a monkey enjoying a day out in the park on roller skates.

  • Prompt

    Describe this in detail.

    Response

    The image shows a half of an avocado on the left side, with a pit visible inside. On the right side, there is a stylized letter "A" in a bold, uppercase font. The avocado and the letter "A" are separated by a white arrow pointing from the avocado to the letter "A," indicating a transformation or relationship between the two. The background is a light beige color.

  • Prompt

    Describe this in detail.

    Response

    The image depicts a dog, specifically a Golden Retriever, enjoying a slice of watermelon. The dog is positioned on the left side of the frame, with its head lowered towards the watermelon. The watermelon is placed on a wooden table, and the dog appears to be biting into the fruit, indicating it is eating it. The background is blurred, suggesting a natural outdoor setting with greenery. The overall scene conveys a sense of relaxation and enjoyment.

03

Acerca de Qwen2-VL 7B Instruct

ResumenA fecha de 23 de septiembre de 2026

Qwen2-VL 7B Instruct es un modelo de Community en la categoría Multimodal. En Railwail, Qwen2-VL 7B Instruct cuesta ≈ 0,0024 US$ por ejecución.

Qwen2-VL 7B Instruct is Alibaba's open-weights vision-language model packaged as a Replicate endpoint. It accepts an image (or video) plus a text prompt and answers questions, describes scenes, reads text in images and extracts structured data. Naive dynamic resolution and M-RoPE let it handle varied aspect ratios and longer visual inputs. A good self-hostable alternative to hosted VLMs for OCR and document tasks.
04

Precios

Precios en dólares estadounidenses. El uso se cobra con créditos prepagados.
Ejecución típica (≈ 2 s en L40S)0,0024 US$ por ejecución
Tiempo de GPU (L40S)0,00117 US$ por segundo de GPU
  • Se factura por el tiempo de GPU que realmente tarda la ejecución. Cuando comienza la ejecución, se reserva 3× el precio típico de tu saldo y se liquida después.
  • 1 crédito = 0,01 US$

Calculadora de costes

Calculadora de precios

s

Típico según el proveedor: aprox. 2,1 s

Total

0,24 US$

24 créditos

Por ejecución

0,0024 US$ · 0,24 créditos

Se factura el tiempo real de GPU; este es un estimado.

05

API

Llama a Qwen2-VL 7B Instruct con tu clave API de Railwail. Usa este ID de modelo en la solicitud:

Sin ejemplo de API verificado

La API pública pasa un formato de entrada diferente al que necesita este modelo. Usa el playground anterior.

06

Especificaciones

ID del modelo
qwen2-vl-7b-instruct-rep
Desarrollador
Community
Categoría
Multimodal
Entrada
Texto, Imagen, Vídeo
Salida
Texto
Facturación
Por uso (tokens o tiempo de GPU)
Entrada del catálogo actualizada
23 de septiembre de 2026

Parámetros de entrada

Entradas y configuración del esquema de entrada del modelo. El ejemplo en la sección API muestra cuáles de ellas acepta la API.

  • promptObligatorio

    Question about the image or video

    Tipo: Texto
    Predeterminado: –
    Valores permitidos: Hasta 16.000 caracteres
  • image_url

    Image or video URL to analyze

    Tipo: Texto
    Predeterminado: –
    Valores permitidos: –
  • max_tokens
    Tipo: Número entero
    Predeterminado: 1024
    Valores permitidos: 1 a 4096
  • temperature
    Tipo: Número
    Predeterminado: 0.7
    Valores permitidos: 0 a 2

Etiquetas

  • replicate
  • qwen
  • alibaba
  • vision-understanding
  • open-weights
  • ocr
07

Preguntas frecuentes

¿Qué es Qwen2-VL 7B Instruct?

Qwen2-VL 7B Instruct es un modelo de Community en la categoría Multimodal.

¿Cuánto cuesta Qwen2-VL 7B Instruct en Railwail?

En Railwail, Qwen2-VL 7B Instruct cuesta ≈ 0,0024 US$ por ejecución. Se te cobra por lo que cada solicitud realmente consume. El uso se paga con créditos prepagados; 1 crédito equivale a 0,01 US$.

¿Qué configuraciones admite Qwen2-VL 7B Instruct?

Según su esquema de entrada, Qwen2-VL 7B Instruct conoce estos parámetros: prompt (Hasta 16.000 caracteres), image_url, max_tokens (1 a 4096) y temperature (0 a 2).

¿Qué velocidad tiene Qwen2-VL 7B Instruct?

Aún no hay suficientes ejecuciones medidas de Qwen2-VL 7B Instruct en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.

¿Es Qwen2-VL 7B Instruct mejor que BLIP?

Depende de la tarea. Qwen2-VL 7B Instruct (Community) y BLIP (Salesforce) son ambos modelos en la categoría Multimodal. La página de comparación muestra sus precios y especificaciones lado a lado.

Comparar Qwen2-VL 7B Instruct y BLIP

¿Puede Qwen2-VL 7B Instruct procesar imágenes?

Sí. Qwen2-VL 7B Instruct acepta imágenes como entrada además de texto.

08

Modelos comparables

Todos en esta categoría
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 US$/ejecución

    88 % más barato por unidad

    Comparar Qwen2-VL 7B Instruct vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    ≈ 0,0457 US$/ejecución

    1804 % más caro por unidad

    Comparar Qwen2-VL 7B Instruct vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    ≈ 0,0050 US$/ejecución

    108 % más caro por unidad

    Comparar Qwen2-VL 7B Instruct vs. Depth Anything v2

Todos los modelos a través de una API

Una clave API para todos los modelos en Railwail. El uso se cobra desde créditos prepagados, 1 crédito = 0,01 US$.