CogVideoX-5B (open)

VideogenereringBeschikbaar
van Zhipu AIModel-ID: cogvideox-5b-open

Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

Prijs
≈ US$ 0,4081/uitvoering
Invoer → Uitvoer
Tekst → Video
Ontwikkelaar
Zhipu AI
Bijgewerkt
23 september 2026
01

Playground

CogVideoX-5B (open) proberen

Geen invoerformulier

≈ US$ 0,4081/uitvoering

Voor dit model is nog geen invoerformulier beschikbaar

De invoeren zijn nog niet gedocumenteerd. Om te voorkomen dat een uitvoering mislukt door onjuiste invoer, bieden we hier geen formulier aan. Kies in plaats daarvan een vergelijkbaar model.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical atmosphere of this unique musical performance.

    Lengte: 0:06

    Settings

    num_frames
    49
  • Prompt

    A suited astronaut, with the red dust of Mars clinging to their boots, reaches out to shake hands with an alien being, their skin a shimmering blue, under the pink-tinged sky of the fourth planet. In the background, a sleek silver rocket, a beacon of human ingenuity, stands tall, its engines powered down, as the two representatives of different worlds exchange a historic greeting amidst the desolate beauty of the Martian landscape.

    Lengte: 0:06

    Settings

    num_frames
    49
  • Prompt

    A garden comes to life as a kaleidoscope of butterflies flutters amidst the blossoms, their delicate wings casting shadows on the petals below. In the background, a grand fountain cascades water with a gentle splendor, its rhythmic sound providing a soothing backdrop. Beneath the cool shade of a mature tree, a solitary wooden chair invites solitude and reflection, its smooth surface worn by the touch of countless visitors seeking a moment of tranquility in nature's embrace.

    Lengte: 0:06

    Settings

    num_frames
    49
03

Over CogVideoX-5B (open)

SamengevatPer 23 september 2026

CogVideoX-5B (open) is een model van Zhipu AI in de categorie Videogenerering. Op Railwail kost CogVideoX-5B (open) ≈ US$ 0,4081 per uitvoering.

Achtergrond

Over THUDM (Tsinghua University KEG Lab) / Zhipu AI

Opgericht 2019 · Beijing, China

THUDM (the Knowledge Engineering Group at Tsinghua University) and its commercial arm Zhipu AI are among China's leading open-source generative-AI research groups. The lab is best known for the GLM language-model family (GLM-130B, ChatGLM) and the CogVLM/CogVideo vision projects. CogVideo was introduced in 2022 as one of the first publicly released 9B-parameter text-to-video transformers, followed by CogVideoX in August 2024, which released 2B and 5B variants under an open-weight licence. Zhipu AI raised more than $1.5B by 2025 and is one of the four 'AI tigers' of China alongside Moonshot, MiniMax and Baichuan. CogVideoX is widely adopted as a research and fine-tuning base because of its permissive licence and reproducible training recipe.

THUDM (Tsinghua University KEG Lab) / Zhipu AI bezoeken

Architectuur

Diffusion Transformer (DiT) with expert Transformer blocks and 3D causal VAE

CogVideoX is a latent video diffusion model built around a 3D causal Variational Autoencoder that compresses video into a compact latent grid across both spatial and temporal axes. On top of this latent space, an Expert Transformer (a diffusion transformer with separate text and video expert streams sharing self-attention) jointly denoises text-conditioned latent video at multiple resolutions. The architecture employs 3D Rotary Position Embeddings (3D-RoPE) for spatio-temporal positions, an Adaptive LayerNorm controlled by the diffusion timestep, and Flow Matching with v-prediction as the training objective. Training proceeds in stages: first low-resolution images, then short low-resolution videos, then high-resolution video up to 720x480 at 8 fps for 6 seconds. The team curated a large filtered video corpus with dense bilingual captions produced by a fine-tuned vision-language model. CogVideoX-5B supports text-to-video and an image-to-video variant (CogVideoX-5B-I2V).

Parameters
5 billion (also 2B variant)
Context
226 tokens

Mogelijkheden

  • Open-weight 5B text-to-video model (Apache 2.0-style permissive licence on weights)
  • Generates 6-second clips at 720x480 / 8 fps natively (~49 frames)
  • Image-to-video variant CogVideoX-5B-I2V for conditioning on a first frame
  • Strong prompt adherence on complex compositions and motion verbs
  • Bilingual English/Chinese prompting
  • Runs on a single 24-48 GB consumer-class GPU with optimisations (CPU offload, INT8)
  • Fine-tunable with LoRA and full-parameter fine-tuning
  • Frequently extended by community for longer durations via temporal tiling
  • Best for: research, open-source pipelines, custom fine-tunes, on-prem video generation.

Training & licentie

Trained on a curated multi-million-clip video corpus with bilingual dense captions generated by a fine-tuned video captioning model. Data is heavily filtered for aesthetic quality, motion coherence and caption alignment. Exact token / clip counts are reported in the paper.

Licentie: Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).

Veiligheidstests: Released with safety filtering on training data and recommended NSFW classifiers for downstream deployment; no formal RSP-style policy.

Bekende beperkingen

  • Maximum native duration ~6 seconds
  • Resolution capped at 720x480 in the 5B base model
  • No audio generation
  • Slower than closed commercial APIs on similar hardware
  • Occasional anatomical artifacts and limb drift on fast motion
04

Prijzen

Prijzen in US-dollars. Het gebruik wordt in rekening gebracht via vooraf gekochte credits.
Typische uitvoering (≈ 349 s op L40S)US$ 0,4081 per uitvoering
GPU-tijd (L40S)US$ 0,00117 per GPU-seconde
  • Afgerekend wordt de GPU-tijd die de run werkelijk gebruikt. Wanneer de run start, wordt het 3-voudige van de typische prijs van uw saldo gereserveerd en later verrekend.
  • 1 credit = US$ 0,01

Kostencalculator

Prijscalculator

s

Typisch volgens de provider: ongeveer 348,7 s

Totaal

US$ 40,81

4.081 credits

Per run

US$ 0,4081 · 40,81 credits

Afgerekend wordt de werkelijke GPU-tijd; dit is een schatting.

05

API

Roep CogVideoX-5B (open) aan met je Railwail API-sleutel. Gebruik deze model-ID in het verzoek:

Geen geverifieerd API-voorbeeld

De invoer van dit model is nog niet gedocumenteerd.

06

Specificaties

Model-ID
cogvideox-5b-open
Ontwikkelaar
Zhipu AI
Invoer
Tekst
Uitvoer
Video
Facturering
Op basis van gebruik (tokens of GPU-tijd)
Modelgrootte
5 billion (also 2B variant)
Licentie
Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).
Catalogusitem bijgewerkt
23 september 2026

Tags

  • zhipu
  • tsinghua
  • cogvideox
  • text-to-video
  • open-weights
  • pricing-tbd
07

Gebruiksscenario's

Waarvoor het wordt gebruikt

  • Open-source video generation pipelines
  • Academic research on video diffusion
  • On-prem creative tooling
  • LoRA fine-tunes for stylised video
  • Image-to-video animation
  • Benchmark baseline for new video models
08

Veelgestelde vragen

Wat is CogVideoX-5B (open)?

CogVideoX-5B (open) is een model van Zhipu AI in de categorie Videogenerering.

Hoeveel kost CogVideoX-5B (open) op Railwail?

Op Railwail kost CogVideoX-5B (open) ≈ US$ 0,4081 per uitvoering. U betaalt voor wat elke aanvraag daadwerkelijk verbruikt. Gebruik wordt betaald met vooraf gekochte credits; 1 credit is gelijk aan US$ 0,01.

Hoe snel is CogVideoX-5B (open)?

Er zijn nog niet genoeg gemeten runs van CogVideoX-5B (open) op Railwail om een uitvoeringstijd op te geven. Dit hangt af van de invoer, de instellingen en de belasting bij de provider.

Is CogVideoX-5B (open) beter dan Google Veo 3.1?

Dat hangt van de taak af. CogVideoX-5B (open) (Zhipu AI) en Google Veo 3.1 (Google DeepMind) zijn beide modellen in de categorie Videogenerering. De vergelijkingspagina toont hun prijzen en specificaties naast elkaar.

CogVideoX-5B (open) en Google Veo 3.1 vergelijken
09

Vergelijkbare modellen

Alle in deze categorie

Alle modellen via één API

Één API-sleutel voor elk model op Railwail. Gebruik wordt afgerekend via vooraf gekochte credits, 1 credit = US$ 0,01.