BLIP
blip-captioningSalesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- Pris
- ≈ 0,00030 US$/kørsel
- Input → output
- Tekst + Billede → Tekst
- Udvikler
- Salesforce
- Opdateret
- 23. september 2026
Playground
Prøv BLIP
Input & resultat
Denne kørsel
ca. 0,0003 US$ · 0,03 credits
0,0009Â US$ (0,09 credits) reserveres ved start; den faktiske GPU-tid faktureres.
Ny her?
10 gratis credits (0,10 US$) når du tilmelder dig med Google
Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits. Nok til 111 kørsler af denne model.
Examples
InputPrompt
image_captioning
Output (JSON, shortened)
[ { "text": "Caption: a woman sitting on the beach with a dog" } ]
InputPrompt
what is in the sky?
Output (JSON, shortened)
[ { "text": "Answer: moon" } ]
InputPrompt
what is in the lake?
Output (JSON, shortened)
[ { "text": "Answer: swans" } ]
Om BLIP
BLIP er en model af Salesforce i kategorien Multimodal. På Railwail koster BLIP ≈ 0,00030 US$ pr. kørsel.
Priser
| Typisk kørsel (≈ 1 s på T4) | 0,00030 US$ pr. kørsel |
|---|---|
| GPU-tid (T4) | 0,00027Â US$ pr. GPU-sekund |
- Faktureres efter den GPU-tid, som kørslen faktisk tager. Når kørslen starter, reserveres 3× den typiske pris fra din saldo og afregnes efterfølgende.
- 1 kredit = 0,01Â US$
Omkostningsberegner
Prisberegner
Typisk ifølge udbyderen: ca. 1 s
I alt
0,03Â US$
3 credits
Pr. kørsel
0,0003 US$ · 0,03 credits
Fakturering efter faktisk GPU-tid; dette er et estimat.
API
Intet bekræftet API-eksempel
Den offentlige API sender et andet inputformat end det, som denne model har brug for. Brug playground ovenfor.
Specifikationer
- Model-ID
blip-captioning- Udvikler
- Salesforce
- Kategori
- Multimodal
- Input
- Tekst, Billede
- Output
- Tekst
- Fakturering
- Efter forbrug (tokens eller GPU-tid)
- Katalogelement opdateret
- 23. september 2026
Inputparametre
Inputs og indstillinger fra modellens inputskema. Eksemplet i API-afsnittet viser, hvilke af dem API'en accepterer.
imagepåkrævetImage to caption or ask about
Type: TekstStandard: –Tilladte værdier: –taskType: ValgStandard:image_captioningTilladte værdier: image_captioning eller visual_question_answeringquestionQuestion for visual question answering mode
Type: TekstStandard: –Tilladte værdier: –
Tags
- replicate
- blip
- captioning
- vqa
- salesforce
- vision-understanding
- image
Ofte stillede spørgsmål
Hvad er BLIP?
BLIP er en model fra Salesforce i kategorien Multimodal.
Hvad koster BLIP på Railwail?
På Railwail koster BLIP ≈ 0,00030 US$ pr. kørsel. Du betaler for det, som hver anmodning faktisk bruger. Forbrug betales fra forudbetalte credits; 1 credit svarer til 0,01 US$.
Hvilke indstillinger understøtter BLIP?
Ifølge dens inputskema kender BLIP disse parametre: image, task (image_captioning eller visual_question_answering) og question.
Hvor hurtig er BLIP?
Der er endnu ikke nok målte kørsler af BLIP på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.
Er BLIP bedre end CLIP Interrogator?
Det afhænger af opgaven. BLIP (Salesforce) og CLIP Interrogator (Community) er begge modeller i kategorien Multimodal. Sammenligningssiden viser deres priser og specifikationer side om side.
Sammenlign BLIP og CLIP InterrogatorKan BLIP behandle billeder?
Ja. BLIP accepterer billeder som input ud over tekst.
Sammenlignelige modeller
Alle i denne kategori- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.
Alle modeller via én API
En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.