CogVideoX-5B (open)

VideogenereringTilgjengelig
av Zhipu AIModell-ID: cogvideox-5b-open

Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

Pris
≈ 0,4081 USD/kjøring
Input → output
Tekst → Video
Utvikler
Zhipu AI
Oppdatert
23. september 2026
01

Playground

Prøv CogVideoX-5B (open)

Ingen innmaskering

≈ 0,4081 USD/kjøring

Ingen innmaskering for denne modellen ennå

Inndataene er ikke dokumentert ennå. For at ingen kjøring skal mislykkes på grunn av feil inndata, tilbyr vi ikke et skjema her. Velg en sammenlignbar modell i stedet.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical atmosphere of this unique musical performance.

    Lengde: 0:06

    Settings

    num_frames
    49
  • Prompt

    A suited astronaut, with the red dust of Mars clinging to their boots, reaches out to shake hands with an alien being, their skin a shimmering blue, under the pink-tinged sky of the fourth planet. In the background, a sleek silver rocket, a beacon of human ingenuity, stands tall, its engines powered down, as the two representatives of different worlds exchange a historic greeting amidst the desolate beauty of the Martian landscape.

    Lengde: 0:06

    Settings

    num_frames
    49
  • Prompt

    A garden comes to life as a kaleidoscope of butterflies flutters amidst the blossoms, their delicate wings casting shadows on the petals below. In the background, a grand fountain cascades water with a gentle splendor, its rhythmic sound providing a soothing backdrop. Beneath the cool shade of a mature tree, a solitary wooden chair invites solitude and reflection, its smooth surface worn by the touch of countless visitors seeking a moment of tranquility in nature's embrace.

    Lengde: 0:06

    Settings

    num_frames
    49
03

Om CogVideoX-5B (open)

Kort sagtPer 23. september 2026

CogVideoX-5B (open) er en modell fra Zhipu AI i kategorien Videogenerering. På Railwail koster CogVideoX-5B (open) ≈ 0,4081 USD per kjøring.

Bakgrunn

Om THUDM (Tsinghua University KEG Lab) / Zhipu AI

Grunnlagt 2019 · Beijing, China

THUDM (the Knowledge Engineering Group at Tsinghua University) and its commercial arm Zhipu AI are among China's leading open-source generative-AI research groups. The lab is best known for the GLM language-model family (GLM-130B, ChatGLM) and the CogVLM/CogVideo vision projects. CogVideo was introduced in 2022 as one of the first publicly released 9B-parameter text-to-video transformers, followed by CogVideoX in August 2024, which released 2B and 5B variants under an open-weight licence. Zhipu AI raised more than $1.5B by 2025 and is one of the four 'AI tigers' of China alongside Moonshot, MiniMax and Baichuan. CogVideoX is widely adopted as a research and fine-tuning base because of its permissive licence and reproducible training recipe.

Besøk THUDM (Tsinghua University KEG Lab) / Zhipu AI

Arkitektur

Diffusion Transformer (DiT) with expert Transformer blocks and 3D causal VAE

CogVideoX is a latent video diffusion model built around a 3D causal Variational Autoencoder that compresses video into a compact latent grid across both spatial and temporal axes. On top of this latent space, an Expert Transformer (a diffusion transformer with separate text and video expert streams sharing self-attention) jointly denoises text-conditioned latent video at multiple resolutions. The architecture employs 3D Rotary Position Embeddings (3D-RoPE) for spatio-temporal positions, an Adaptive LayerNorm controlled by the diffusion timestep, and Flow Matching with v-prediction as the training objective. Training proceeds in stages: first low-resolution images, then short low-resolution videos, then high-resolution video up to 720x480 at 8 fps for 6 seconds. The team curated a large filtered video corpus with dense bilingual captions produced by a fine-tuned vision-language model. CogVideoX-5B supports text-to-video and an image-to-video variant (CogVideoX-5B-I2V).

Parametere
5 billion (also 2B variant)
Kontekst
226 tokens

Evner

  • Open-weight 5B text-to-video model (Apache 2.0-style permissive licence on weights)
  • Generates 6-second clips at 720x480 / 8 fps natively (~49 frames)
  • Image-to-video variant CogVideoX-5B-I2V for conditioning on a first frame
  • Strong prompt adherence on complex compositions and motion verbs
  • Bilingual English/Chinese prompting
  • Runs on a single 24-48 GB consumer-class GPU with optimisations (CPU offload, INT8)
  • Fine-tunable with LoRA and full-parameter fine-tuning
  • Frequently extended by community for longer durations via temporal tiling
  • Best for: research, open-source pipelines, custom fine-tunes, on-prem video generation.

Trening og lisens

Trained on a curated multi-million-clip video corpus with bilingual dense captions generated by a fine-tuned video captioning model. Data is heavily filtered for aesthetic quality, motion coherence and caption alignment. Exact token / clip counts are reported in the paper.

Lisens: Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).

Sikkerhetstesting: Released with safety filtering on training data and recommended NSFW classifiers for downstream deployment; no formal RSP-style policy.

Kjente begrensninger

  • Maximum native duration ~6 seconds
  • Resolution capped at 720x480 in the 5B base model
  • No audio generation
  • Slower than closed commercial APIs on similar hardware
  • Occasional anatomical artifacts and limb drift on fast motion
04

Priser

Priser i amerikanske dollar. Bruk belastes fra forhåndsbetalt kreditt.
Typisk kjøring (≈ 349 s på L40S)0,4081 USD per kjøring
GPU-tid (L40S)0,00117 USD per GPU-sekund
  • Faktureres etter GPU-tiden kjøringen faktisk tar. Når kjøringen starter, blir 3× den typiske prisen reservert fra saldoen din og gjort opp etterpå.
  • 1 kreditt = 0,01 USD

Kostnadsberegner

Prisberegner

s

Typisk ifølge leverandøren: ca. 348,7 s

Totalt

40,81 USD

4 081 credits

Per kjøring

0,4081 USD · 40,81 credits

Fakturert etter faktisk GPU-tid; dette er et estimat.

05

API

Ring CogVideoX-5B (open) med din Railwail API-nøkkel. Bruk denne modell-IDen i forespørselen:

Ingen verifisert API-eksempel

Inngangene til denne modellen er ikke dokumentert ennå.

06

Spesifikasjoner

Modell-ID
cogvideox-5b-open
Utvikler
Zhipu AI
Inndata
Tekst
Utdata
Video
Fakturering
Etter bruk (tokens eller GPU-tid)
Modellstørrelse
5 billion (also 2B variant)
Lisens
Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).
Katalogoppføring oppdatert
23. september 2026

Merkelapper

  • zhipu
  • tsinghua
  • cogvideox
  • text-to-video
  • open-weights
  • pricing-tbd
07

Brukstilfeller

Hva det brukes til

  • Open-source video generation pipelines
  • Academic research on video diffusion
  • On-prem creative tooling
  • LoRA fine-tunes for stylised video
  • Image-to-video animation
  • Benchmark baseline for new video models
08

Ofte stilte spørsmål

Hva er CogVideoX-5B (open)?

CogVideoX-5B (open) er en modell fra Zhipu AI i kategorien Videogenerering.

Hvor mye koster CogVideoX-5B (open) på Railwail?

På Railwail koster CogVideoX-5B (open) ≈ 0,4081 USD per kjøring. Du betaler for det hver forespørsel faktisk bruker. Bruk betales fra forhåndskjøpte kreditter; 1 kreditt tilsvarer 0,01 USD.

Hvor rask er CogVideoX-5B (open)?

Det finnes ennå ikke nok målte kjøringer av CogVideoX-5B (open) på Railwail til å angi en kjøretid. Det avhenger av inndataene, innstillingene og belastningen hos leverandøren.

Er CogVideoX-5B (open) bedre enn Google Veo 3.1?

Det avhenger av oppgaven. CogVideoX-5B (open) (Zhipu AI) og Google Veo 3.1 (Google DeepMind) er begge modeller i kategorien Videogenerering. Sammenligningssiden viser prisene og spesifikasjonene deres side ved side.

Sammenlign CogVideoX-5B (open) og Google Veo 3.1
09

Sammenlignbare modeller

Alle i denne kategorien

Alle modeller gjennom én API

Én API-nøkkel for alle modeller på Railwail. Bruk belastes fra forhåndsbetalt kreditt, 1 kreditt = 0,01 USD.