Yi-VL 34B

MultimodalUnavailable
by 01.AIModel ID: yi-vl-34b

01.AI Yi-VL 34B vision-language model. Bilingual (CN/EN) image understanding, strong CMMMU and MMMU performance among open-weights VLMs.

Status
Unavailable
Context
4,096 tokens
Max. output
2,048 tokens
Input → output
Text + Image → Text
Developer
01.AI
Updated
September 23, 2026

Yi-VL 34B is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ US$0.00030/run

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    US$6.00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    US$3.60/1M in

02

Playground

Try Yi-VL 34B

No input form

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    Describe this image.

    Response

    In the heart of a bustling city, a man in a vibrant yellow shirt and blue pants is engaged in an unusual activity. He's perched atop a yellow taxi cab, his attention focused on a pile of neatly folded blue shirts. The taxi, a common sight in the city, is parked on the side of a busy street. The man appears to be in the process of either packing or unpacking these shirts, adding a touch of human interest to the urban landscape. The street itself is a hive of activity. Cars of various colors are parked along the curb, their metallic bodies gleaming under the city lights. Buildings tower in the background, their windows reflecting the hustle and bustle of the street below. Adding to the urban charm of the scene are pink flags fluttering from the buildings. They hang in the air, their bright color contrasting with the more muted tones of the cityscape. Their presence suggests a festive or special occasion, adding a sense of celebration to the everyday scene. Overall, this image captures a moment of unexpected human activity amidst the routine of city life, hinting at stories and events that lie beneath the surface of the urban landscape.

04

About Yi-VL 34B

TL;DRAs of September 23, 2026

Yi-VL 34B is a model by 01.AI in the Multimodal category. Yi-VL 34B is currently not available on Railwail. The context window holds 4,096 tokens, and one response can be up to 2,048 tokens long.

05

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

06

API

Call Yi-VL 34B with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

07

Specifications

Model ID
yi-vl-34b
Developer
01.AI
Category
Multimodal
Input
Text, Image
Output
Text
Context window
4,096 tokens
Max. output
2,048 tokens
Catalog entry updated
September 23, 2026

Tags

  • replicate
  • multimodal
  • vision-understanding
  • 01ai
  • open-weights
  • bilingual
08

Frequently asked questions

What is Yi-VL 34B?

Yi-VL 34B is a model by 01.AI in the Multimodal category. It is listed on Railwail but cannot be run at the moment.

How much does Yi-VL 34B cost on Railwail?

Yi-VL 34B cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

What is the context window of Yi-VL 34B?

The context window of Yi-VL 34B holds 4,096 tokens. One response can be up to 2,048 tokens long.

How fast is Yi-VL 34B?

There are not enough measured runs of Yi-VL 34B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Yi-VL 34B better than BLIP?

That depends on the task. Yi-VL 34B (01.AI) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Yi-VL 34B and BLIP

Can Yi-VL 34B process images?

Yes. Yi-VL 34B accepts images as input in addition to text.

Can I use Yi-VL 34B right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.