DeepSeek-VL 7B
deepseek-vl-7bDeepSeek-VL 7B chat model. Vision-language model with hybrid vision encoder and strong real-world visual question answering performance.
- Preț
- ≈ 0,0086 USD/rulare
- Context
- 4.096 tokeni
- Ieșire max.
- 2.048 tokeni
- Intrare → ieșire
- Text + Imagine → Text
- Dezvoltator
- DeepSeek
- Actualizat
- 23 septembrie 2026
Playground
Încearcă DeepSeek-VL 7B
Fără formular de intrare
Nicio formă de intrare pentru acest model încă
Intrările sale nu sunt documentate încă. Pentru ca nicio rulare să nu eșueze din cauza unei intrări greșite, nu oferim un formular aici. Alege în schimb un model comparabil.
Examples
InputPrompt
Describe each stage of this image
Response
The image depicts a three-stage process for training a vision-language model. 1. Stage 1: Training VL Adapter: In this stage, a vision-language adapter is trained using supervised fine-tuning. The adapter is trained on image-text pairs and pure language sequences. 2. Stage 2: Joint VL Pre-training: In this stage, a joint vision-language model is pre-trained using self-supervised learning. The model is trained on image-text pairs and pure language sequences. 3. Stage 3: Supervised Fine-tuning: In this stage, the model is fine-tuned on supervised tasks using image-text pairs and pure language sequences. The model is trained using a hybrid vision-language adapter, which combines a vision-language adapter with a language model. The model is trained on a variety of tasks, including image captioning, visual question answering, and visual reasoning. The model is able to understand the visual content of an image and generate a natural language description or answer.
InputPrompt
Describe this image
Response
The image you've provided is a meme featuring a French Bulldog lying on its side on a blue wooden floor. The dog's head is resting on its paws, and it appears to be sleeping. The meme has a text overlay that reads "MONDAY. JUST...MONDAY." This suggests that the meme is making a humorous comment about the start of the work week. The dog's relaxed posture and the text overlay together create a light-hearted and relatable message.
Despre DeepSeek-VL 7B
DeepSeek-VL 7B este un model de DeepSeek din categoria Multimodal. Pe Railwail, DeepSeek-VL 7B costă ≈ 0,0086 USD per rulare. Fereastra de context conține 4.096 token-uri, iar un răspuns poate fi lung de până la 2.048 token-uri.
Prețuri
| Rulare tipică (≈ 7 s pe L40S) | 0,0086 USD per rulare |
|---|---|
| Timp GPU (L40S) | 0,00117 USD per secundă GPU |
- Se facturează timpul GPU pe care îl necesită efectiv rularea. La pornire, se rezervă din soldul tău 3× prețul tipic și se decontează ulterior.
- 1 credit = 0,01 USD
Calculator de costuri
Calculator de preț
Tipic conform furnizorului: aprox. 7,3 s
Total
0,86 USD
86 credite
Pe rulare
0,0086 USD · 0,86 credite
Se factură după timpul GPU real; aceasta este o estimare.
API
Niciun exemplu API verificat
API-ul public transmite un format de intrare diferit de ceea ce are nevoie acest model. Folosește playground-ul de mai sus.
Specificații
- ID model
deepseek-vl-7b- Dezvoltator
- DeepSeek
- Categorie
- Multimodal
- Intrare
- Text, Imagine
- Ieșire
- Text
- Fereastră de context
- 4.096 tokeni
- Ieșire max.
- 2.048 tokeni
- Facturare
- După utilizare (tokeni sau timp GPU)
- Intrare catalog actualizată
- 23 septembrie 2026
Etichete
- replicate
- multimodal
- vision-understanding
- deepseek
- open-weights
Întrebări frecvente
Ce este DeepSeek-VL 7B?
DeepSeek-VL 7B este un model de DeepSeek din categoria Multimodal.
Cât costă DeepSeek-VL 7B pe Railwail?
Pe Railwail, DeepSeek-VL 7B costă ≈ 0,0086 USD per rulare. Ți se percepe taxa pentru ceea ce fiecare cerere folosește efectiv. Utilizarea se plătește din credite prepay; 1 credit egal cu 0,01 USD.
Care este fereastra de context a DeepSeek-VL 7B?
Fereastra de context a DeepSeek-VL 7B conține 4.096 token-uri. Un răspuns poate fi lung de până la 2.048 token-uri.
Cât de rapid este DeepSeek-VL 7B?
Nu sunt suficiente rulări măsurate ale DeepSeek-VL 7B pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.
Este DeepSeek-VL 7B mai bun decât BLIP?
Depinde de sarcină. DeepSeek-VL 7B (DeepSeek) și BLIP (Salesforce) sunt ambele modele din categoria Multimodal. Pagina de comparație arată prețurile și specificațiile lor una lângă alta.
Compară DeepSeek-VL 7B și BLIPPoate DeepSeek-VL 7B procesa imagini?
Da. DeepSeek-VL 7B acceptă imagini ca intrare, pe lângă text.
Modele comparabile
Toate din această categorie- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
Toate modelele printr-o singură API
O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.