Qwen2-VL 7B Instruct
qwen2-vl-7b-instruct-repAlibaba Qwen2-VL 7B served on Replicate. Open-weights vision-language model that chats about images and video, with dynamic resolution and strong OCR and document QA for its size.
- Price
- โ US$0.0024/run
- Input โ output
- Text + Image + Video โ Text
- Developer
- Community
- Updated
- 23 September 2026
Playground
Try Qwen2-VL 7B Instruct
Input & output
This run
about US$0.0024 ยท 0.24 credits
US$0.0072 (0.72 credits) are reserved at the start; the actual GPU time is billed.
New here?
10 free credits (US$0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 13 runs of this model.
Examples
Prompt
Describe this video in detail.
Response
The video features a monkey riding a skateboard in a park. The monkey is wearing roller skates and is seen moving forward on the skateboard, with its arms outstretched for balance. The park is filled with people sitting on benches and walking around, enjoying the sunny day. The monkey appears to be enjoying itself, and the scene is filled with a sense of fun and playfulness. The park is surrounded by trees and greenery, creating a peaceful and relaxing atmosphere. Overall, the video captures a unique and entertaining moment of a monkey enjoying a day out in the park on roller skates.
Prompt
Describe this in detail.
Response
The image shows a half of an avocado on the left side, with a pit visible inside. On the right side, there is a stylized letter "A" in a bold, uppercase font. The avocado and the letter "A" are separated by a white arrow pointing from the avocado to the letter "A," indicating a transformation or relationship between the two. The background is a light beige color.
Prompt
Describe this in detail.
Response
The image depicts a dog, specifically a Golden Retriever, enjoying a slice of watermelon. The dog is positioned on the left side of the frame, with its head lowered towards the watermelon. The watermelon is placed on a wooden table, and the dog appears to be biting into the fruit, indicating it is eating it. The background is blurred, suggesting a natural outdoor setting with greenery. The overall scene conveys a sense of relaxation and enjoyment.
About Qwen2-VL 7B Instruct
Qwen2-VL 7B Instruct is a model by Community in the Multimodal category. On Railwail, Qwen2-VL 7B Instruct costs โ US$0.0024 per run.
Pricing
| Typical run (โ 2 s on L40S) | US$0.0024 per run |
|---|---|
| GPU time (L40S) | US$0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = US$0.01
Cost calculator
Price calculator
Typical according to the provider: about 2.1 s
Total
US$0.24
24 credits
Per run
US$0.0024 ยท 0.24 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
qwen2-vl-7b-instruct-rep- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image, Video
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- 23 September 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
promptrequiredQuestion about the image or video
Type: TextDefault: โAllowed values: up to 16,000 charactersimage_urlImage or video URL to analyze
Type: TextDefault: โAllowed values: โmax_tokensType: IntegerDefault:1024Allowed values: 1 to 4,096temperatureType: NumberDefault:0.7Allowed values: 0 to 2
Tags
- replicate
- qwen
- alibaba
- vision-understanding
- open-weights
- ocr
Frequently asked questions
What is Qwen2-VL 7B Instruct?
Qwen2-VL 7B Instruct is a model by Community in the Multimodal category.
How much does Qwen2-VL 7B Instruct cost on Railwail?
On Railwail, Qwen2-VL 7B Instruct costs โ US$0.0024 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.
Which settings does Qwen2-VL 7B Instruct support?
According to its input schema, Qwen2-VL 7B Instruct knows these parameters: prompt (up to 16,000 characters), image_url, max_tokens (1 to 4,096) and temperature (0 to 2).
How fast is Qwen2-VL 7B Instruct?
There are not enough measured runs of Qwen2-VL 7B Instruct on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Qwen2-VL 7B Instruct better than BLIP?
That depends on the task. Qwen2-VL 7B Instruct (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Qwen2-VL 7B Instruct and BLIPCan Qwen2-VL 7B Instruct process images?
Yes. Qwen2-VL 7B Instruct accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.