CogVideoX-5B (open)

Generación de vídeosDisponible
de Zhipu AIID del modelo: cogvideox-5b-open

Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

Precio
≈ 0,4081 US$/ejecución
Entrada → Salida
Texto → Vídeo
Desarrollador
Zhipu AI
Actualizado
23 de septiembre de 2026
01

Playground

Probar CogVideoX-5B (open)

Sin formulario de entrada

≈ 0,4081 US$/ejecución

Aún no hay formulario de entrada para este modelo

Sus entradas aún no están documentadas. Para que ninguna ejecución falle por una entrada incorrecta, no ofrecemos un formulario aquí. Elige un modelo comparable en su lugar.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical atmosphere of this unique musical performance.

    Duración: 0:06

    Settings

    num_frames
    49
  • Prompt

    A suited astronaut, with the red dust of Mars clinging to their boots, reaches out to shake hands with an alien being, their skin a shimmering blue, under the pink-tinged sky of the fourth planet. In the background, a sleek silver rocket, a beacon of human ingenuity, stands tall, its engines powered down, as the two representatives of different worlds exchange a historic greeting amidst the desolate beauty of the Martian landscape.

    Duración: 0:06

    Settings

    num_frames
    49
  • Prompt

    A garden comes to life as a kaleidoscope of butterflies flutters amidst the blossoms, their delicate wings casting shadows on the petals below. In the background, a grand fountain cascades water with a gentle splendor, its rhythmic sound providing a soothing backdrop. Beneath the cool shade of a mature tree, a solitary wooden chair invites solitude and reflection, its smooth surface worn by the touch of countless visitors seeking a moment of tranquility in nature's embrace.

    Duración: 0:06

    Settings

    num_frames
    49
03

Acerca de CogVideoX-5B (open)

ResumenA fecha de 23 de septiembre de 2026

CogVideoX-5B (open) es un modelo de Zhipu AI en la categoría Generación de vídeos. En Railwail, CogVideoX-5B (open) cuesta ≈ 0,4081 US$ por ejecución.

Fondo

Acerca de THUDM (Tsinghua University KEG Lab) / Zhipu AI

Fundado en 2019 · Beijing, China

THUDM (the Knowledge Engineering Group at Tsinghua University) and its commercial arm Zhipu AI are among China's leading open-source generative-AI research groups. The lab is best known for the GLM language-model family (GLM-130B, ChatGLM) and the CogVLM/CogVideo vision projects. CogVideo was introduced in 2022 as one of the first publicly released 9B-parameter text-to-video transformers, followed by CogVideoX in August 2024, which released 2B and 5B variants under an open-weight licence. Zhipu AI raised more than $1.5B by 2025 and is one of the four 'AI tigers' of China alongside Moonshot, MiniMax and Baichuan. CogVideoX is widely adopted as a research and fine-tuning base because of its permissive licence and reproducible training recipe.

Visitar THUDM (Tsinghua University KEG Lab) / Zhipu AI

Arquitectura

Diffusion Transformer (DiT) with expert Transformer blocks and 3D causal VAE

CogVideoX is a latent video diffusion model built around a 3D causal Variational Autoencoder that compresses video into a compact latent grid across both spatial and temporal axes. On top of this latent space, an Expert Transformer (a diffusion transformer with separate text and video expert streams sharing self-attention) jointly denoises text-conditioned latent video at multiple resolutions. The architecture employs 3D Rotary Position Embeddings (3D-RoPE) for spatio-temporal positions, an Adaptive LayerNorm controlled by the diffusion timestep, and Flow Matching with v-prediction as the training objective. Training proceeds in stages: first low-resolution images, then short low-resolution videos, then high-resolution video up to 720x480 at 8 fps for 6 seconds. The team curated a large filtered video corpus with dense bilingual captions produced by a fine-tuned vision-language model. CogVideoX-5B supports text-to-video and an image-to-video variant (CogVideoX-5B-I2V).

Parámetros
5 billion (also 2B variant)
Contexto
226 tokens

Capacidades

  • Open-weight 5B text-to-video model (Apache 2.0-style permissive licence on weights)
  • Generates 6-second clips at 720x480 / 8 fps natively (~49 frames)
  • Image-to-video variant CogVideoX-5B-I2V for conditioning on a first frame
  • Strong prompt adherence on complex compositions and motion verbs
  • Bilingual English/Chinese prompting
  • Runs on a single 24-48 GB consumer-class GPU with optimisations (CPU offload, INT8)
  • Fine-tunable with LoRA and full-parameter fine-tuning
  • Frequently extended by community for longer durations via temporal tiling
  • Best for: research, open-source pipelines, custom fine-tunes, on-prem video generation.

Entrenamiento y licencia

Trained on a curated multi-million-clip video corpus with bilingual dense captions generated by a fine-tuned video captioning model. Data is heavily filtered for aesthetic quality, motion coherence and caption alignment. Exact token / clip counts are reported in the paper.

Licencia: Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).

Pruebas de seguridad: Released with safety filtering on training data and recommended NSFW classifiers for downstream deployment; no formal RSP-style policy.

Limitaciones conocidas

  • Maximum native duration ~6 seconds
  • Resolution capped at 720x480 in the 5B base model
  • No audio generation
  • Slower than closed commercial APIs on similar hardware
  • Occasional anatomical artifacts and limb drift on fast motion
04

Precios

Precios en dólares estadounidenses. El uso se cobra con créditos prepagados.
Ejecución típica (≈ 349 s en L40S)0,4081 US$ por ejecución
Tiempo de GPU (L40S)0,00117 US$ por segundo de GPU
  • Se factura por el tiempo de GPU que realmente tarda la ejecución. Cuando comienza la ejecución, se reserva 3× el precio típico de tu saldo y se liquida después.
  • 1 crédito = 0,01 US$

Calculadora de costes

Calculadora de precios

s

Típico según el proveedor: aprox. 348,7 s

Total

40,81 US$

4081 créditos

Por ejecución

0,4081 US$ · 40,81 créditos

Se factura el tiempo real de GPU; este es un estimado.

05

API

Llama a CogVideoX-5B (open) con tu clave API de Railwail. Usa este ID de modelo en la solicitud:

Sin ejemplo de API verificado

Las entradas de este modelo aún no están documentadas.

06

Especificaciones

ID del modelo
cogvideox-5b-open
Desarrollador
Zhipu AI
Entrada
Texto
Salida
Vídeo
Facturación
Por uso (tokens o tiempo de GPU)
Tamaño del modelo
5 billion (also 2B variant)
Licencia
Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).
Entrada del catálogo actualizada
23 de septiembre de 2026

Etiquetas

  • zhipu
  • tsinghua
  • cogvideox
  • text-to-video
  • open-weights
  • pricing-tbd
07

Casos de uso

Para qué se utiliza

  • Open-source video generation pipelines
  • Academic research on video diffusion
  • On-prem creative tooling
  • LoRA fine-tunes for stylised video
  • Image-to-video animation
  • Benchmark baseline for new video models
08

Preguntas frecuentes

¿Qué es CogVideoX-5B (open)?

CogVideoX-5B (open) es un modelo de Zhipu AI en la categoría Generación de vídeos.

¿Cuánto cuesta CogVideoX-5B (open) en Railwail?

En Railwail, CogVideoX-5B (open) cuesta ≈ 0,4081 US$ por ejecución. Se te cobra por lo que cada solicitud realmente consume. El uso se paga con créditos prepagados; 1 crédito equivale a 0,01 US$.

¿Qué velocidad tiene CogVideoX-5B (open)?

Aún no hay suficientes ejecuciones medidas de CogVideoX-5B (open) en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.

¿Es CogVideoX-5B (open) mejor que Google Veo 3.1?

Depende de la tarea. CogVideoX-5B (open) (Zhipu AI) y Google Veo 3.1 (Google DeepMind) son ambos modelos en la categoría Generación de vídeos. La página de comparación muestra sus precios y especificaciones lado a lado.

Comparar CogVideoX-5B (open) y Google Veo 3.1
09

Modelos comparables

Todos en esta categoría

Todos los modelos a través de una API

Una clave API para todos los modelos en Railwail. El uso se cobra desde créditos prepagados, 1 crédito = 0,01 US$.