DeepSeek-VL 7B
deepseek-vl-7bDeepSeek-VL 7B chat model. Vision-language model with hybrid vision encoder and strong real-world visual question answering performance.
- Price
- โ $0.0086/run
- Context
- 4,096 tokens
- Max. output
- 2,048 tokens
- Input โ output
- Text + Image โ Text
- Developer
- DeepSeek
- Updated
- September 23, 2026
Playground
Try DeepSeek-VL 7B
No input form
No input form for this model yet
Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.
Examples
InputPrompt
Describe each stage of this image
Response
The image depicts a three-stage process for training a vision-language model. 1. Stage 1: Training VL Adapter: In this stage, a vision-language adapter is trained using supervised fine-tuning. The adapter is trained on image-text pairs and pure language sequences. 2. Stage 2: Joint VL Pre-training: In this stage, a joint vision-language model is pre-trained using self-supervised learning. The model is trained on image-text pairs and pure language sequences. 3. Stage 3: Supervised Fine-tuning: In this stage, the model is fine-tuned on supervised tasks using image-text pairs and pure language sequences. The model is trained using a hybrid vision-language adapter, which combines a vision-language adapter with a language model. The model is trained on a variety of tasks, including image captioning, visual question answering, and visual reasoning. The model is able to understand the visual content of an image and generate a natural language description or answer.
InputPrompt
Describe this image
Response
The image you've provided is a meme featuring a French Bulldog lying on its side on a blue wooden floor. The dog's head is resting on its paws, and it appears to be sleeping. The meme has a text overlay that reads "MONDAY. JUST...MONDAY." This suggests that the meme is making a humorous comment about the start of the work week. The dog's relaxed posture and the text overlay together create a light-hearted and relatable message.
About DeepSeek-VL 7B
DeepSeek-VL 7B is a model by DeepSeek in the Multimodal category. On Railwail, DeepSeek-VL 7B costs โ $0.0086 per run. The context window holds 4,096 tokens, and one response can be up to 2,048 tokens long.
Pricing
| Typical run (โ 7 s on L40S) | $0.0086 per run |
|---|---|
| GPU time (L40S) | $0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 7.3 s
Total
$0.86
86 credits
Per run
$0.0086 ยท 0.86 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
deepseek-vl-7b- Developer
- DeepSeek
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Context window
- 4,096 tokens
- Max. output
- 2,048 tokens
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Tags
- replicate
- multimodal
- vision-understanding
- deepseek
- open-weights
Frequently asked questions
What is DeepSeek-VL 7B?
DeepSeek-VL 7B is a model by DeepSeek in the Multimodal category.
How much does DeepSeek-VL 7B cost on Railwail?
On Railwail, DeepSeek-VL 7B costs โ $0.0086 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
What is the context window of DeepSeek-VL 7B?
The context window of DeepSeek-VL 7B holds 4,096 tokens. One response can be up to 2,048 tokens long.
How fast is DeepSeek-VL 7B?
There are not enough measured runs of DeepSeek-VL 7B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is DeepSeek-VL 7B better than BLIP?
That depends on the task. DeepSeek-VL 7B (DeepSeek) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare DeepSeek-VL 7B and BLIPCan DeepSeek-VL 7B process images?
Yes. DeepSeek-VL 7B accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.