BLIP

MultimodalTilgængelig
af SalesforceModell-ID: blip-captioning

Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

Pris
≈ 0,00030 US$/kørsel
Input → output
Tekst + Billede → Tekst
Udvikler
Salesforce
Opdateret
23. september 2026
01

Playground

Prøv BLIP

Input & resultat

≈ 0,00030 US$/kørsel
Prøv BLIP

Question for visual question answering mode

Resultat
Svaret vises her.

Denne kørsel

ca. 0,0003 US$ · 0,03 credits

0,0009 US$ (0,09 credits) reserveres ved start; den faktiske GPU-tid faktureres.

Ny her?

10 gratis credits (0,10 US$) når du tilmelder dig med Google

Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits. Nok til 111 kørsler af denne model.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    image_captioning

    Output (JSON, shortened)

    [
      {
        "text": "Caption: a woman sitting on the beach with a dog"
      }
    ]
  • InputInput

    Prompt

    what is in the sky?

    Output (JSON, shortened)

    [
      {
        "text": "Answer: moon"
      }
    ]
  • InputInput

    Prompt

    what is in the lake?

    Output (JSON, shortened)

    [
      {
        "text": "Answer: swans"
      }
    ]
03

Om BLIP

Kort sagtFra 23. september 2026

BLIP er en model af Salesforce i kategorien Multimodal. På Railwail koster BLIP ≈ 0,00030 US$ pr. kørsel.

BLIP (Bootstrapping Language-Image Pre-training) from Salesforce Research unifies captioning and VQA in one model. In caption mode it returns a concise description of the scene; in question-answering mode it answers a free-form question grounded in the image. It is one of the most-run captioning models on Replicate and a common building block for image search and accessibility alt-text.
04

Priser

Priser i US-dollar. Forbrug debiteres fra forudbetalte credits.
Typisk kørsel (≈ 1 s på T4)0,00030 US$ pr. kørsel
GPU-tid (T4)0,00027 US$ pr. GPU-sekund
  • Faktureres efter den GPU-tid, som kørslen faktisk tager. NÃ¥r kørslen starter, reserveres 3× den typiske pris fra din saldo og afregnes efterfølgende.
  • 1 kredit = 0,01 US$

Omkostningsberegner

Prisberegner

s

Typisk ifølge udbyderen: ca. 1 s

I alt

0,03 US$

3 credits

Pr. kørsel

0,0003 US$ · 0,03 credits

Fakturering efter faktisk GPU-tid; dette er et estimat.

05

API

Kald BLIP med din Railwail API-nøgle. Brug dette model-ID i anmodningen:

Intet bekræftet API-eksempel

Den offentlige API sender et andet inputformat end det, som denne model har brug for. Brug playground ovenfor.

06

Specifikationer

Model-ID
blip-captioning
Udvikler
Salesforce
Kategori
Multimodal
Input
Tekst, Billede
Output
Tekst
Fakturering
Efter forbrug (tokens eller GPU-tid)
Katalogelement opdateret
23. september 2026

Inputparametre

Inputs og indstillinger fra modellens inputskema. Eksemplet i API-afsnittet viser, hvilke af dem API'en accepterer.

  • imagepÃ¥krævet

    Image to caption or ask about

    Type: Tekst
    Standard: –
    Tilladte værdier: –
  • task
    Type: Valg
    Standard: image_captioning
    Tilladte værdier: image_captioning eller visual_question_answering
  • question

    Question for visual question answering mode

    Type: Tekst
    Standard: –
    Tilladte værdier: –

Tags

  • replicate
  • blip
  • captioning
  • vqa
  • salesforce
  • vision-understanding
  • image
07

Ofte stillede spørgsmål

Hvad er BLIP?

BLIP er en model fra Salesforce i kategorien Multimodal.

Hvad koster BLIP på Railwail?

På Railwail koster BLIP ≈ 0,00030 US$ pr. kørsel. Du betaler for det, som hver anmodning faktisk bruger. Forbrug betales fra forudbetalte credits; 1 credit svarer til 0,01 US$.

Hvilke indstillinger understøtter BLIP?

Ifølge dens inputskema kender BLIP disse parametre: image, task (image_captioning eller visual_question_answering) og question.

Hvor hurtig er BLIP?

Der er endnu ikke nok målte kørsler af BLIP på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.

Er BLIP bedre end CLIP Interrogator?

Det afhænger af opgaven. BLIP (Salesforce) og CLIP Interrogator (Community) er begge modeller i kategorien Multimodal. Sammenligningssiden viser deres priser og specifikationer side om side.

Sammenlign BLIP og CLIP Interrogator

Kan BLIP behandle billeder?

Ja. BLIP accepterer billeder som input ud over tekst.

08

Sammenlignelige modeller

Alle i denne kategori
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    ≈ 0,0457 US$/kørsel

    15.133 % dyrere pr. enhed

    Sammenlign BLIP og CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    ≈ 0,0050 US$/kørsel

    1.567 % dyrere pr. enhed

    Sammenlign BLIP og Depth Anything v2
  • Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.

    ≈ 0,018 US$/kørsel

    5.900 % dyrere pr. enhed

    Sammenlign BLIP og SAM 2 (Segment Anything 2)

Alle modeller via én API

En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.