CogVideoX-5B (open)

VideogenerierungVerfügbar
von Zhipu AIModell-ID: cogvideox-5b-open

Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

Preis
ca. $ 0,4081/Lauf
Eingabe → Ausgabe
Text → Video
Entwickler
Zhipu AI
Aktualisiert
23. September 2026
01

Playground

CogVideoX-5B (open) ausprobieren

Keine Eingabemaske

ca. $ 0,4081/Lauf

Für dieses Modell ist noch keine Eingabemaske hinterlegt

Seine Eingaben sind noch nicht erfasst. Damit kein Lauf an einer falschen Eingabe scheitert, bieten wir hier kein Formular an. Wähle stattdessen ein vergleichbares Modell.

02

Beispiele

Echte Ergebnisse aus den öffentlichen Beispielen dieses Modells auf Replicate, mit dem Prompt und den Einstellungen, die sie erzeugt haben. Sie wurden nicht live auf dieser Seite erzeugt.
  • Prompt

    A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical atmosphere of this unique musical performance.

    Länge: 0:06

    Einstellungen

    num_frames
    49
  • Prompt

    A suited astronaut, with the red dust of Mars clinging to their boots, reaches out to shake hands with an alien being, their skin a shimmering blue, under the pink-tinged sky of the fourth planet. In the background, a sleek silver rocket, a beacon of human ingenuity, stands tall, its engines powered down, as the two representatives of different worlds exchange a historic greeting amidst the desolate beauty of the Martian landscape.

    Länge: 0:06

    Einstellungen

    num_frames
    49
  • Prompt

    A garden comes to life as a kaleidoscope of butterflies flutters amidst the blossoms, their delicate wings casting shadows on the petals below. In the background, a grand fountain cascades water with a gentle splendor, its rhythmic sound providing a soothing backdrop. Beneath the cool shade of a mature tree, a solitary wooden chair invites solitude and reflection, its smooth surface worn by the touch of countless visitors seeking a moment of tranquility in nature's embrace.

    Länge: 0:06

    Einstellungen

    num_frames
    49
03

Über CogVideoX-5B (open)

Kurz gesagtStand: 23. September 2026

CogVideoX-5B (open) ist ein Modell von Zhipu AI aus der Kategorie Videogenerierung. Über Railwail kostet CogVideoX-5B (open) ca. $ 0,4081 pro Lauf.

Hintergrund

Über THUDM (Tsinghua University KEG Lab) / Zhipu AI

Gegründet 2019 · Beijing, China

THUDM (the Knowledge Engineering Group at Tsinghua University) and its commercial arm Zhipu AI are among China's leading open-source generative-AI research groups. The lab is best known for the GLM language-model family (GLM-130B, ChatGLM) and the CogVLM/CogVideo vision projects. CogVideo was introduced in 2022 as one of the first publicly released 9B-parameter text-to-video transformers, followed by CogVideoX in August 2024, which released 2B and 5B variants under an open-weight licence. Zhipu AI raised more than $1.5B by 2025 and is one of the four 'AI tigers' of China alongside Moonshot, MiniMax and Baichuan. CogVideoX is widely adopted as a research and fine-tuning base because of its permissive licence and reproducible training recipe.

THUDM (Tsinghua University KEG Lab) / Zhipu AI besuchen

Architektur

Diffusion Transformer (DiT) with expert Transformer blocks and 3D causal VAE

CogVideoX is a latent video diffusion model built around a 3D causal Variational Autoencoder that compresses video into a compact latent grid across both spatial and temporal axes. On top of this latent space, an Expert Transformer (a diffusion transformer with separate text and video expert streams sharing self-attention) jointly denoises text-conditioned latent video at multiple resolutions. The architecture employs 3D Rotary Position Embeddings (3D-RoPE) for spatio-temporal positions, an Adaptive LayerNorm controlled by the diffusion timestep, and Flow Matching with v-prediction as the training objective. Training proceeds in stages: first low-resolution images, then short low-resolution videos, then high-resolution video up to 720x480 at 8 fps for 6 seconds. The team curated a large filtered video corpus with dense bilingual captions produced by a fine-tuned vision-language model. CogVideoX-5B supports text-to-video and an image-to-video variant (CogVideoX-5B-I2V).

Parameter
5 billion (also 2B variant)
Kontext
226 Token

Funktionen

  • Open-weight 5B text-to-video model (Apache 2.0-style permissive licence on weights)
  • Generates 6-second clips at 720x480 / 8 fps natively (~49 frames)
  • Image-to-video variant CogVideoX-5B-I2V for conditioning on a first frame
  • Strong prompt adherence on complex compositions and motion verbs
  • Bilingual English/Chinese prompting
  • Runs on a single 24-48 GB consumer-class GPU with optimisations (CPU offload, INT8)
  • Fine-tunable with LoRA and full-parameter fine-tuning
  • Frequently extended by community for longer durations via temporal tiling
  • Best for: research, open-source pipelines, custom fine-tunes, on-prem video generation.

Training & Lizenz

Trained on a curated multi-million-clip video corpus with bilingual dense captions generated by a fine-tuned video captioning model. Data is heavily filtered for aesthetic quality, motion coherence and caption alignment. Exact token / clip counts are reported in the paper.

Lizenz: Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).

Sicherheitstests: Released with safety filtering on training data and recommended NSFW classifiers for downstream deployment; no formal RSP-style policy.

Bekannte Einschränkungen

  • Maximum native duration ~6 seconds
  • Resolution capped at 720x480 in the 5B base model
  • No audio generation
  • Slower than closed commercial APIs on similar hardware
  • Occasional anatomical artifacts and limb drift on fast motion
04

Preise

Preise in US-Dollar. Abgerechnet wird über vorab gekaufte Credits.
Typischer Lauf (ca. 349 s auf L40S)$ 0,4081 pro Lauf
GPU-Zeit (L40S)$ 0,00117 pro GPU-Sekunde
  • Abgerechnet wird die GPU-Zeit, die der Lauf tatsächlich braucht. Beim Start wird das 3-Fache des typischen Preises vom Guthaben vorgemerkt und danach verrechnet.
  • 1 Credit = $ 0,01

Kostenrechner

Preisrechner

s

Typisch laut Anbieter: ca. 348,7 s

Gesamt

$ 40,81

4 081 Credits

Je Lauf

$ 0,4081 · 40,81 Credits

Abgerechnet wird die tatsächliche GPU-Zeit; der Wert ist eine Schätzung.

05

API

Rufe CogVideoX-5B (open) mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Kein geprüftes API-Beispiel

Die Eingaben dieses Modells sind noch nicht erfasst.

06

Spezifikationen

Modell-ID
cogvideox-5b-open
Entwickler
Zhipu AI
Eingabe
Text
Ausgabe
Video
Abrechnung
Nach Verbrauch (Token bzw. GPU-Zeit)
Modellgröße
5 billion (also 2B variant)
Lizenz
Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).
Katalogeintrag aktualisiert
23. September 2026

Schlagwörter

  • zhipu
  • tsinghua
  • cogvideox
  • text-to-video
  • open-weights
  • pricing-tbd
07

Einsatzgebiete

Wofür es genutzt wird

  • Open-source video generation pipelines
  • Academic research on video diffusion
  • On-prem creative tooling
  • LoRA fine-tunes for stylised video
  • Image-to-video animation
  • Benchmark baseline for new video models
08

Häufige Fragen

Was ist CogVideoX-5B (open)?

CogVideoX-5B (open) ist ein Modell von Zhipu AI aus der Kategorie Videogenerierung.

Was kostet CogVideoX-5B (open) bei Railwail?

Über Railwail kostet CogVideoX-5B (open) ca. $ 0,4081 pro Lauf. Abgerechnet wird, was jede Anfrage tatsächlich verbraucht. Bezahlt wird mit vorab gekauften Credits; 1 Credit entspricht $ 0,01.

Wie schnell ist CogVideoX-5B (open)?

Für CogVideoX-5B (open) gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist CogVideoX-5B (open) besser als Google Veo 3.1?

Das hängt von der Aufgabe ab. CogVideoX-5B (open) (Zhipu AI) und Google Veo 3.1 (Google DeepMind) sind beide Modelle aus der Kategorie Videogenerierung. Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

CogVideoX-5B (open) und Google Veo 3.1 vergleichen
09

Vergleichbare Modelle

Alle dieser Kategorie

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0,01.