LLaVA-OneVision 72B

MultimodalUnavailable
by ReplicateModel ID: llava-onevision-72b

LMMs-Lab LLaVA-OneVision 72B. Unified single-image, multi-image and video instruction-tuned VLM with task-transfer across modalities.

Status
Unavailable
Context
32,768 tokens
Max. output
4,096 tokens
Input β†’ output
Text + Image + Video β†’ Text
Developer
Replicate
Updated
June 25, 2026

LLaVA-OneVision 72B is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    β‰ˆ $0.00030/run

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    $6.00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    $3.60/1M in

02

Playground

Try LLaVA-OneVision 72B

Input & output

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try LLaVA-OneVision 72B

0 / 16,000

User message or question about the image

Optional image URL to analyze

Advanced settings (4)

0 / 8,000

Optional system instruction

Output
The answer appears here.

This run

No price – currently unavailable.

New here?

5 free credits ($0.05) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About LLaVA-OneVision 72B

TL;DRAs of June 25, 2026

LLaVA-OneVision 72B is a model by Replicate in the Multimodal category. LLaVA-OneVision 72B is currently not available on Railwail. The context window holds 32,768 tokens, and one response can be up to 4,096 tokens long.

04

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

05

API

Call LLaVA-OneVision 72B with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
llava-onevision-72b
Developer
Replicate
Category
Multimodal
Input
Text, Image, Video
Output
Text
Context window
32,768 tokens
Max. output
4,096 tokens
Catalog entry updated
June 25, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • promptrequired

    User message or question about the image

    Type: Text
    Default: –
    Allowed values: up to 16,000 characters
  • top_p
    Type: Number
    Default: 1
    Allowed values: 0 to 1
  • stream
    Type: Yes/no
    Default: false
    Allowed values: –
  • image_url

    Optional image URL to analyze

    Type: Text
    Default: –
    Allowed values: –
  • max_tokens
    Type: Integer
    Default: 1024
    Allowed values: 1 to 4,096
  • temperature
    Type: Number
    Default: 0.7
    Allowed values: 0 to 2
  • system_prompt

    Optional system instruction

    Type: Text
    Default: –
    Allowed values: up to 8,000 characters

Tags

  • replicate
  • multimodal
  • vision-understanding
  • llava
  • open-weights
07

Frequently asked questions

What is LLaVA-OneVision 72B?

LLaVA-OneVision 72B is a model by Replicate in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does LLaVA-OneVision 72B cost on Railwail?

LLaVA-OneVision 72B cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of LLaVA-OneVision 72B?

The context window of LLaVA-OneVision 72B holds 32,768 tokens. One response can be up to 4,096 tokens long.

How fast is LLaVA-OneVision 72B?

There are not enough measured runs of LLaVA-OneVision 72B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is LLaVA-OneVision 72B better than BLIP?

That depends on the task. LLaVA-OneVision 72B (Replicate) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare LLaVA-OneVision 72B and BLIP

Can LLaVA-OneVision 72B process images?

Yes. LLaVA-OneVision 72B accepts images as input in addition to text.

Can I use LLaVA-OneVision 72B right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.