CogVLM2 19B

MultimodalAvailable
by CommunityModel ID: cogvlm2-19b

Tsinghua CogVLM2 19B with Llama-3 8B base plus 11B vision expert. Strong document understanding and visual reasoning, 8k context.

Price
โ‰ˆ US$0.0114/run
Context
8,192 tokens
Max. output
2,048 tokens
Input โ†’ output
Text + Image โ†’ Text
Developer
Community
Updated
23 September 2026
01

Playground

Try CogVLM2 19B

No input form

โ‰ˆ US$0.0114/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    Describe this image.

    Response

    The image captures a well-lit, modern library or bookstore with a distinct industrial aesthetic. The main focus is a large, wooden bookshelf filled with an assortment of books, creating a warm and inviting atmosphere. The bookshelf is positioned against a rustic brick wall, which adds a touch of vintage charm to the space. The room is illuminated by hanging light bulbs, which dangle from the ceiling in a casual manner. There are also decorative elements such as a potted plant, a small framed sign, and a table with various items on it, enhancing the cozy ambiance. A person is seated at a desk in the foreground, suggesting the space is functional for reading or studying.

03

About CogVLM2 19B

TL;DRAs of 23 September 2026

CogVLM2 19B is a model by Community in the Multimodal category. On Railwail, CogVLM2 19B costs โ‰ˆ US$0.0114 per run. The context window holds 8,192 tokens, and one response can be up to 2,048 tokens long.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 10 s on L40S)US$0.0114 per run
GPU time (L40S)US$0.00117 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = US$0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 9.7 s

Total

US$1.14

114 credits

Per run

US$0.0114 ยท 1.14 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call CogVLM2 19B with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
cogvlm2-19b
Developer
Community
Category
Multimodal
Input
Text, Image
Output
Text
Context window
8,192 tokens
Max. output
2,048 tokens
Billing
By usage (tokens or GPU time)
Catalog entry updated
23 September 2026

Tags

  • replicate
  • multimodal
  • vision-understanding
  • tsinghua
  • open-weights
07

Frequently asked questions

What is CogVLM2 19B?

CogVLM2 19B is a model by Community in the Multimodal category.

How much does CogVLM2 19B cost on Railwail?

On Railwail, CogVLM2 19B costs โ‰ˆ US$0.0114 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.

What is the context window of CogVLM2 19B?

The context window of CogVLM2 19B holds 8,192 tokens. One response can be up to 2,048 tokens long.

How fast is CogVLM2 19B?

There are not enough measured runs of CogVLM2 19B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is CogVLM2 19B better than BLIP?

That depends on the task. CogVLM2 19B (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare CogVLM2 19B and BLIP

Can CogVLM2 19B process images?

Yes. CogVLM2 19B accepts images as input in addition to text.

08

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    โ‰ˆ US$0.00030/run

    97 % cheaper per unit

    Compare CogVLM2 19B vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    โ‰ˆ US$0.0457/run

    301 % more expensive per unit

    Compare CogVLM2 19B vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    โ‰ˆ US$0.0050/run

    56 % cheaper per unit

    Compare CogVLM2 19B vs. Depth Anything v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.