DeepSeek-VL 7B

MultimodalAvailable
by DeepSeekModel ID: deepseek-vl-7b

DeepSeek-VL 7B chat model. Vision-language model with hybrid vision encoder and strong real-world visual question answering performance.

Price
โ‰ˆ $0.0086/run
Context
4,096 tokens
Max. output
2,048 tokens
Input โ†’ output
Text + Image โ†’ Text
Developer
DeepSeek
Updated
September 23, 2026
01

Playground

Try DeepSeek-VL 7B

No input form

โ‰ˆ $0.0086/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    Describe each stage of this image

    Response

    The image depicts a three-stage process for training a vision-language model. 1. Stage 1: Training VL Adapter: In this stage, a vision-language adapter is trained using supervised fine-tuning. The adapter is trained on image-text pairs and pure language sequences. 2. Stage 2: Joint VL Pre-training: In this stage, a joint vision-language model is pre-trained using self-supervised learning. The model is trained on image-text pairs and pure language sequences. 3. Stage 3: Supervised Fine-tuning: In this stage, the model is fine-tuned on supervised tasks using image-text pairs and pure language sequences. The model is trained using a hybrid vision-language adapter, which combines a vision-language adapter with a language model. The model is trained on a variety of tasks, including image captioning, visual question answering, and visual reasoning. The model is able to understand the visual content of an image and generate a natural language description or answer.

  • InputInput

    Prompt

    Describe this image

    Response

    The image you've provided is a meme featuring a French Bulldog lying on its side on a blue wooden floor. The dog's head is resting on its paws, and it appears to be sleeping. The meme has a text overlay that reads "MONDAY. JUST...MONDAY." This suggests that the meme is making a humorous comment about the start of the work week. The dog's relaxed posture and the text overlay together create a light-hearted and relatable message.

03

About DeepSeek-VL 7B

TL;DRAs of September 23, 2026

DeepSeek-VL 7B is a model by DeepSeek in the Multimodal category. On Railwail, DeepSeek-VL 7B costs โ‰ˆ $0.0086 per run. The context window holds 4,096 tokens, and one response can be up to 2,048 tokens long.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 7 s on L40S)$0.0086 per run
GPU time (L40S)$0.00117 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 7.3 s

Total

$0.86

86 credits

Per run

$0.0086 ยท 0.86 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call DeepSeek-VL 7B with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
deepseek-vl-7b
Developer
DeepSeek
Category
Multimodal
Input
Text, Image
Output
Text
Context window
4,096 tokens
Max. output
2,048 tokens
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Tags

  • replicate
  • multimodal
  • vision-understanding
  • deepseek
  • open-weights
07

Frequently asked questions

What is DeepSeek-VL 7B?

DeepSeek-VL 7B is a model by DeepSeek in the Multimodal category.

How much does DeepSeek-VL 7B cost on Railwail?

On Railwail, DeepSeek-VL 7B costs โ‰ˆ $0.0086 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

What is the context window of DeepSeek-VL 7B?

The context window of DeepSeek-VL 7B holds 4,096 tokens. One response can be up to 2,048 tokens long.

How fast is DeepSeek-VL 7B?

There are not enough measured runs of DeepSeek-VL 7B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is DeepSeek-VL 7B better than BLIP?

That depends on the task. DeepSeek-VL 7B (DeepSeek) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare DeepSeek-VL 7B and BLIP

Can DeepSeek-VL 7B process images?

Yes. DeepSeek-VL 7B accepts images as input in addition to text.

08

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    โ‰ˆ $0.00030/run

    97 % cheaper per unit

    Compare DeepSeek-VL 7B vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    โ‰ˆ $0.0457/run

    431 % more expensive per unit

    Compare DeepSeek-VL 7B vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    โ‰ˆ $0.0050/run

    42 % cheaper per unit

    Compare DeepSeek-VL 7B vs. Depth Anything v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.