MiniCPM-V 2.6

MultimodalAvailable
by CommunityModel ID: minicpm-v-2-6

OpenBMB MiniCPM-V 2.6. 8B vision-language model with strong single-image, multi-image and video understanding plus OCR capabilities.

Price
≈ US$0.0015/run
Context
32,768 tokens
Max. output
4,096 tokens
Input → output
Text + Image + Video → Text
Developer
Community
Updated
September 23, 2026
01

Playground

Try MiniCPM-V 2.6

No input form

≈ US$0.0015/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Response

    视频展示了几个场景,强调了学习和实践的重要性。首先,一个穿着围裙的人正在用蓝色胶带做标记,暗示着一个装修或艺术项目。接着,一个小组在办公室环境中开会,强调了团队合作和知识分享。然后,视频展示了一个人在篮球场上投篮,象征着通过实践学习和成长。最后,展示了一个人在厨房里撒糖粉,可能是在烘焙,强调了将知识应用于实践。这些场景共同传达了不断学习和实践知识的重要性,无论是在个人成长还是职业发展中。

  • Response

    这幅图片展示了一个年轻男性的动画角色,他有着深棕色的头发和棕色的眼睛,微微泛红的脸颊,表明他可能感到害羞或兴奋。他穿着一件简单的白色短袖衬衫,背着一个黑色背包,暗示他可能是一个学生或旅行者。他的姿势轻松,一只手托着头,另一只手放在膝盖上,表明他处于放松或思考的状态。背景是一个郁郁葱葱的森林,阳光透过树叶洒下,营造出宁静的氛围。柔和的光线和散落的心形光斑增添了场景的梦幻和浪漫氛围。

03

About MiniCPM-V 2.6

TL;DRAs of September 23, 2026

MiniCPM-V 2.6 is a model by Community in the Multimodal category. On Railwail, MiniCPM-V 2.6 costs ≈ US$0.0015 per run. The context window holds 32,768 tokens, and one response can be up to 4,096 tokens long.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (≈ 1 s on L40S)US$0.0015 per run
GPU time (L40S)US$0.00117 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3× the typical price is reserved from your balance and settled afterwards.
  • 1 credit = US$0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 1.2 s

Total

US$0.15

15 credits

Per run

US$0.0015 · 0.15 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call MiniCPM-V 2.6 with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
minicpm-v-2-6
Developer
Community
Category
Multimodal
Input
Text, Image, Video
Output
Text
Context window
32,768 tokens
Max. output
4,096 tokens
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Tags

  • replicate
  • multimodal
  • vision-understanding
  • open-weights
  • small
07

Frequently asked questions

What is MiniCPM-V 2.6?

MiniCPM-V 2.6 is a model by Community in the Multimodal category.

How much does MiniCPM-V 2.6 cost on Railwail?

On Railwail, MiniCPM-V 2.6 costs ≈ US$0.0015 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.

What is the context window of MiniCPM-V 2.6?

The context window of MiniCPM-V 2.6 holds 32,768 tokens. One response can be up to 4,096 tokens long.

How fast is MiniCPM-V 2.6?

There are not enough measured runs of MiniCPM-V 2.6 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is MiniCPM-V 2.6 better than BLIP?

That depends on the task. MiniCPM-V 2.6 (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare MiniCPM-V 2.6 and BLIP

Can MiniCPM-V 2.6 process images?

Yes. MiniCPM-V 2.6 accepts images as input in addition to text.

08

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ US$0.00030/run

    80 % cheaper per unit

    Compare MiniCPM-V 2.6 vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    ≈ US$0.0457/run

    2,947 % more expensive per unit

    Compare MiniCPM-V 2.6 vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    ≈ US$0.0050/run

    233 % more expensive per unit

    Compare MiniCPM-V 2.6 vs. Depth Anything v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.