BLIP

MultimodalAvailable
by SalesforceModel ID: blip-captioning

Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

Price
โ‰ˆ US$0.00030/run
Input โ†’ output
Text + Image โ†’ Text
Developer
Salesforce
Updated
September 23, 2026
01

Playground

Try BLIP

Input & output

โ‰ˆ US$0.00030/run
Try BLIP

Question for visual question answering mode

Image to caption or ask about

Output
The answer appears here.

This run

about US$0.0003 ยท 0.03 credits

US$0.0009 (0.09 credits) are reserved at the start; the actual GPU time is billed.

New here?

10 free credits (US$0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 111 runs of this model.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    image_captioning

    Output (JSON, shortened)

    [
      {
        "text": "Caption: a woman sitting on the beach with a dog"
      }
    ]
  • InputInput

    Prompt

    what is in the sky?

    Output (JSON, shortened)

    [
      {
        "text": "Answer: moon"
      }
    ]
  • InputInput

    Prompt

    what is in the lake?

    Output (JSON, shortened)

    [
      {
        "text": "Answer: swans"
      }
    ]
03

About BLIP

TL;DRAs of September 23, 2026

BLIP is a model by Salesforce in the Multimodal category. On Railwail, BLIP costs โ‰ˆ US$0.00030 per run.

BLIP (Bootstrapping Language-Image Pre-training) from Salesforce Research unifies captioning and VQA in one model. In caption mode it returns a concise description of the scene; in question-answering mode it answers a free-form question grounded in the image. It is one of the most-run captioning models on Replicate and a common building block for image search and accessibility alt-text.
04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 1 s on T4)US$0.00030 per run
GPU time (T4)US$0.00027 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = US$0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 1 s

Total

US$0.03

3 credits

Per run

US$0.0003 ยท 0.03 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call BLIP with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
blip-captioning
Developer
Salesforce
Category
Multimodal
Input
Text, Image
Output
Text
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • imagerequired

    Image to caption or ask about

    Type: Text
    Default: โ€“
    Allowed values: โ€“
  • task
    Type: Choice
    Default: image_captioning
    Allowed values: image_captioning or visual_question_answering
  • question

    Question for visual question answering mode

    Type: Text
    Default: โ€“
    Allowed values: โ€“

Tags

  • replicate
  • blip
  • captioning
  • vqa
  • salesforce
  • vision-understanding
  • image
07

Frequently asked questions

What is BLIP?

BLIP is a model by Salesforce in the Multimodal category.

How much does BLIP cost on Railwail?

On Railwail, BLIP costs โ‰ˆ US$0.00030 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.

Which settings does BLIP support?

According to its input schema, BLIP knows these parameters: image, task (image_captioning or visual_question_answering) and question.

How fast is BLIP?

There are not enough measured runs of BLIP on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is BLIP better than CLIP Interrogator?

That depends on the task. BLIP (Salesforce) and CLIP Interrogator (Community) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare BLIP and CLIP Interrogator

Can BLIP process images?

Yes. BLIP accepts images as input in addition to text.

08

Comparable models

All in this category
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    โ‰ˆ US$0.0457/run

    15,133 % more expensive per unit

    Compare BLIP vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    โ‰ˆ US$0.0050/run

    1,567 % more expensive per unit

    Compare BLIP vs. Depth Anything v2
  • Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.

    โ‰ˆ US$0.018/run

    5,900 % more expensive per unit

    Compare BLIP vs. SAM 2 (Segment Anything 2)

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.