Yi-VL 34B
yi-vl-34b01.AI Yi-VL 34B vision-language model. Bilingual (CN/EN) image understanding, strong CMMMU and MMMU performance among open-weights VLMs.
- Status
- Unavailable
- Context
- 4,096 tokens
- Max. output
- 2,048 tokens
- Input → output
- Text + Image → Text
- Developer
- 01.AI
- Updated
- September 23, 2026
Yi-VL 34B is currently unavailable
Currently unavailable: this model has been deactivated.
You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.
Go to alternativesComparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
≈ $0.00030/run
- Claude Opus 4.7Anthropic
Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.
$6.00/1M in
- Claude Sonnet 4.6Anthropic
Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.
$3.60/1M in
Playground
Try Yi-VL 34B
No input form
Currently unavailable: this model has been deactivated.
The playground is disabled. You can find comparable models in the same category: Browse alternatives
Examples
InputPrompt
Describe this image.
Response
In the heart of a bustling city, a man in a vibrant yellow shirt and blue pants is engaged in an unusual activity. He's perched atop a yellow taxi cab, his attention focused on a pile of neatly folded blue shirts. The taxi, a common sight in the city, is parked on the side of a busy street. The man appears to be in the process of either packing or unpacking these shirts, adding a touch of human interest to the urban landscape. The street itself is a hive of activity. Cars of various colors are parked along the curb, their metallic bodies gleaming under the city lights. Buildings tower in the background, their windows reflecting the hustle and bustle of the street below. Adding to the urban charm of the scene are pink flags fluttering from the buildings. They hang in the air, their bright color contrasting with the more muted tones of the cityscape. Their presence suggests a festive or special occasion, adding a sense of celebration to the everyday scene. Overall, this image captures a moment of unexpected human activity amidst the routine of city life, hinting at stories and events that lie beneath the surface of the urban landscape.
About Yi-VL 34B
Yi-VL 34B is a model by 01.AI in the Multimodal category. Yi-VL 34B is currently not available on Railwail. The context window holds 4,096 tokens, and one response can be up to 2,048 tokens long.
Pricing
Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
yi-vl-34b- Developer
- 01.AI
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Context window
- 4,096 tokens
- Max. output
- 2,048 tokens
- Catalog entry updated
- September 23, 2026
Tags
- replicate
- multimodal
- vision-understanding
- 01ai
- open-weights
- bilingual
Frequently asked questions
What is Yi-VL 34B?
Yi-VL 34B is a model by 01.AI in the Multimodal category. It is listed on Railwail but cannot be run at the moment.
How much does Yi-VL 34B cost on Railwail?
Yi-VL 34B cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.
What is the context window of Yi-VL 34B?
The context window of Yi-VL 34B holds 4,096 tokens. One response can be up to 2,048 tokens long.
How fast is Yi-VL 34B?
There are not enough measured runs of Yi-VL 34B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Yi-VL 34B better than BLIP?
That depends on the task. Yi-VL 34B (01.AI) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Yi-VL 34B and BLIPCan Yi-VL 34B process images?
Yes. Yi-VL 34B accepts images as input in addition to text.
Can I use Yi-VL 34B right now?
Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.