CogVideoX-5B (open)

Video generationAvailable
by Zhipu AIModel ID: cogvideox-5b-open

Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

Price
โ‰ˆ US$0.4081/run
Input โ†’ output
Text โ†’ Video
Developer
Zhipu AI
Updated
September 23, 2026
01

Playground

Try CogVideoX-5B (open)

No input form

โ‰ˆ US$0.4081/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical atmosphere of this unique musical performance.

    Length: 0:06

    Settings

    num_frames
    49
  • Prompt

    A suited astronaut, with the red dust of Mars clinging to their boots, reaches out to shake hands with an alien being, their skin a shimmering blue, under the pink-tinged sky of the fourth planet. In the background, a sleek silver rocket, a beacon of human ingenuity, stands tall, its engines powered down, as the two representatives of different worlds exchange a historic greeting amidst the desolate beauty of the Martian landscape.

    Length: 0:06

    Settings

    num_frames
    49
  • Prompt

    A garden comes to life as a kaleidoscope of butterflies flutters amidst the blossoms, their delicate wings casting shadows on the petals below. In the background, a grand fountain cascades water with a gentle splendor, its rhythmic sound providing a soothing backdrop. Beneath the cool shade of a mature tree, a solitary wooden chair invites solitude and reflection, its smooth surface worn by the touch of countless visitors seeking a moment of tranquility in nature's embrace.

    Length: 0:06

    Settings

    num_frames
    49
03

About CogVideoX-5B (open)

TL;DRAs of September 23, 2026

CogVideoX-5B (open) is a model by Zhipu AI in the Video generation category. On Railwail, CogVideoX-5B (open) costs โ‰ˆ US$0.4081 per run.

Background

About THUDM (Tsinghua University KEG Lab) / Zhipu AI

Founded 2019 ยท Beijing, China

THUDM (the Knowledge Engineering Group at Tsinghua University) and its commercial arm Zhipu AI are among China's leading open-source generative-AI research groups. The lab is best known for the GLM language-model family (GLM-130B, ChatGLM) and the CogVLM/CogVideo vision projects. CogVideo was introduced in 2022 as one of the first publicly released 9B-parameter text-to-video transformers, followed by CogVideoX in August 2024, which released 2B and 5B variants under an open-weight licence. Zhipu AI raised more than $1.5B by 2025 and is one of the four 'AI tigers' of China alongside Moonshot, MiniMax and Baichuan. CogVideoX is widely adopted as a research and fine-tuning base because of its permissive licence and reproducible training recipe.

Visit THUDM (Tsinghua University KEG Lab) / Zhipu AI

Architecture

Diffusion Transformer (DiT) with expert Transformer blocks and 3D causal VAE

CogVideoX is a latent video diffusion model built around a 3D causal Variational Autoencoder that compresses video into a compact latent grid across both spatial and temporal axes. On top of this latent space, an Expert Transformer (a diffusion transformer with separate text and video expert streams sharing self-attention) jointly denoises text-conditioned latent video at multiple resolutions. The architecture employs 3D Rotary Position Embeddings (3D-RoPE) for spatio-temporal positions, an Adaptive LayerNorm controlled by the diffusion timestep, and Flow Matching with v-prediction as the training objective. Training proceeds in stages: first low-resolution images, then short low-resolution videos, then high-resolution video up to 720x480 at 8 fps for 6 seconds. The team curated a large filtered video corpus with dense bilingual captions produced by a fine-tuned vision-language model. CogVideoX-5B supports text-to-video and an image-to-video variant (CogVideoX-5B-I2V).

Parameters
5 billion (also 2B variant)
Context
226 tokens

Capabilities

  • Open-weight 5B text-to-video model (Apache 2.0-style permissive licence on weights)
  • Generates 6-second clips at 720x480 / 8 fps natively (~49 frames)
  • Image-to-video variant CogVideoX-5B-I2V for conditioning on a first frame
  • Strong prompt adherence on complex compositions and motion verbs
  • Bilingual English/Chinese prompting
  • Runs on a single 24-48 GB consumer-class GPU with optimisations (CPU offload, INT8)
  • Fine-tunable with LoRA and full-parameter fine-tuning
  • Frequently extended by community for longer durations via temporal tiling
  • Best for: research, open-source pipelines, custom fine-tunes, on-prem video generation.

Training & license

Trained on a curated multi-million-clip video corpus with bilingual dense captions generated by a fine-tuned video captioning model. Data is heavily filtered for aesthetic quality, motion coherence and caption alignment. Exact token / clip counts are reported in the paper.

License: Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).

Safety testing: Released with safety filtering on training data and recommended NSFW classifiers for downstream deployment; no formal RSP-style policy.

Known limitations

  • Maximum native duration ~6 seconds
  • Resolution capped at 720x480 in the 5B base model
  • No audio generation
  • Slower than closed commercial APIs on similar hardware
  • Occasional anatomical artifacts and limb drift on fast motion
04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 349 s on L40S)US$0.4081 per run
GPU time (L40S)US$0.00117 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = US$0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 348.7 s

Total

US$40.81

4,081 credits

Per run

US$0.4081 ยท 40.81 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call CogVideoX-5B (open) with your Railwail API key. Use this model ID in the request:

No verified API example

The inputs of this model are not documented yet.

06

Specifications

Model ID
cogvideox-5b-open
Developer
Zhipu AI
Input
Text
Output
Video
Billing
By usage (tokens or GPU time)
Model size
5 billion (also 2B variant)
License
Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).
Catalog entry updated
September 23, 2026

Tags

  • zhipu
  • tsinghua
  • cogvideox
  • text-to-video
  • open-weights
  • pricing-tbd
07

Use cases

What it is used for

  • Open-source video generation pipelines
  • Academic research on video diffusion
  • On-prem creative tooling
  • LoRA fine-tunes for stylised video
  • Image-to-video animation
  • Benchmark baseline for new video models
08

Frequently asked questions

What is CogVideoX-5B (open)?

CogVideoX-5B (open) is a model by Zhipu AI in the Video generation category.

How much does CogVideoX-5B (open) cost on Railwail?

On Railwail, CogVideoX-5B (open) costs โ‰ˆ US$0.4081 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.

How fast is CogVideoX-5B (open)?

There are not enough measured runs of CogVideoX-5B (open) on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is CogVideoX-5B (open) better than Google Veo 3.1?

That depends on the task. CogVideoX-5B (open) (Zhipu AI) and Google Veo 3.1 (Google DeepMind) are both models in the Video generation category. The comparison page shows their prices and specifications side by side.

Compare CogVideoX-5B (open) and Google Veo 3.1
09

Comparable models

All in this category

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.