CogVideoX-5B (open)

Geração de vídeosDisponível
por Zhipu AIID do modelo: cogvideox-5b-open

Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

Preço
≈ US$ 0,4081/execução
Entrada → Saída
Texto → Vídeo
Desenvolvedor
Zhipu AI
Atualizado
23 de setembro de 2026
01

Playground

Experimentar CogVideoX-5B (open)

Sem formulário de entrada

≈ US$ 0,4081/execução

Ainda não há formulário de entrada para este modelo

Suas entradas ainda não foram documentadas. Para que nenhuma execução falhe com uma entrada incorreta, não oferecemos um formulário aqui. Escolha um modelo comparável em vez disso.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical atmosphere of this unique musical performance.

    Duração: 0:06

    Settings

    num_frames
    49
  • Prompt

    A suited astronaut, with the red dust of Mars clinging to their boots, reaches out to shake hands with an alien being, their skin a shimmering blue, under the pink-tinged sky of the fourth planet. In the background, a sleek silver rocket, a beacon of human ingenuity, stands tall, its engines powered down, as the two representatives of different worlds exchange a historic greeting amidst the desolate beauty of the Martian landscape.

    Duração: 0:06

    Settings

    num_frames
    49
  • Prompt

    A garden comes to life as a kaleidoscope of butterflies flutters amidst the blossoms, their delicate wings casting shadows on the petals below. In the background, a grand fountain cascades water with a gentle splendor, its rhythmic sound providing a soothing backdrop. Beneath the cool shade of a mature tree, a solitary wooden chair invites solitude and reflection, its smooth surface worn by the touch of countless visitors seeking a moment of tranquility in nature's embrace.

    Duração: 0:06

    Settings

    num_frames
    49
03

Sobre CogVideoX-5B (open)

ResumoA partir de 23 de setembro de 2026

CogVideoX-5B (open) é um modelo de Zhipu AI na categoria Geração de vídeos. No Railwail, CogVideoX-5B (open) custa ≈ US$ 0,4081 por execução.

Fundo

Sobre THUDM (Tsinghua University KEG Lab) / Zhipu AI

Fundado em 2019 · Beijing, China

THUDM (the Knowledge Engineering Group at Tsinghua University) and its commercial arm Zhipu AI are among China's leading open-source generative-AI research groups. The lab is best known for the GLM language-model family (GLM-130B, ChatGLM) and the CogVLM/CogVideo vision projects. CogVideo was introduced in 2022 as one of the first publicly released 9B-parameter text-to-video transformers, followed by CogVideoX in August 2024, which released 2B and 5B variants under an open-weight licence. Zhipu AI raised more than $1.5B by 2025 and is one of the four 'AI tigers' of China alongside Moonshot, MiniMax and Baichuan. CogVideoX is widely adopted as a research and fine-tuning base because of its permissive licence and reproducible training recipe.

Visite THUDM (Tsinghua University KEG Lab) / Zhipu AI

Arquitetura

Diffusion Transformer (DiT) with expert Transformer blocks and 3D causal VAE

CogVideoX is a latent video diffusion model built around a 3D causal Variational Autoencoder that compresses video into a compact latent grid across both spatial and temporal axes. On top of this latent space, an Expert Transformer (a diffusion transformer with separate text and video expert streams sharing self-attention) jointly denoises text-conditioned latent video at multiple resolutions. The architecture employs 3D Rotary Position Embeddings (3D-RoPE) for spatio-temporal positions, an Adaptive LayerNorm controlled by the diffusion timestep, and Flow Matching with v-prediction as the training objective. Training proceeds in stages: first low-resolution images, then short low-resolution videos, then high-resolution video up to 720x480 at 8 fps for 6 seconds. The team curated a large filtered video corpus with dense bilingual captions produced by a fine-tuned vision-language model. CogVideoX-5B supports text-to-video and an image-to-video variant (CogVideoX-5B-I2V).

Parâmetros
5 billion (also 2B variant)
Contexto
226 tokens

Capacidades

  • Open-weight 5B text-to-video model (Apache 2.0-style permissive licence on weights)
  • Generates 6-second clips at 720x480 / 8 fps natively (~49 frames)
  • Image-to-video variant CogVideoX-5B-I2V for conditioning on a first frame
  • Strong prompt adherence on complex compositions and motion verbs
  • Bilingual English/Chinese prompting
  • Runs on a single 24-48 GB consumer-class GPU with optimisations (CPU offload, INT8)
  • Fine-tunable with LoRA and full-parameter fine-tuning
  • Frequently extended by community for longer durations via temporal tiling
  • Best for: research, open-source pipelines, custom fine-tunes, on-prem video generation.

Treinamento & licença

Trained on a curated multi-million-clip video corpus with bilingual dense captions generated by a fine-tuned video captioning model. Data is heavily filtered for aesthetic quality, motion coherence and caption alignment. Exact token / clip counts are reported in the paper.

Licença: Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).

Testes de segurança: Released with safety filtering on training data and recommended NSFW classifiers for downstream deployment; no formal RSP-style policy.

Limitações conhecidas

  • Maximum native duration ~6 seconds
  • Resolution capped at 720x480 in the 5B base model
  • No audio generation
  • Slower than closed commercial APIs on similar hardware
  • Occasional anatomical artifacts and limb drift on fast motion
04

Preços

Preços em dólares americanos. O uso é cobrado a partir de créditos pré-pagos.
Execução típica (≈ 349 s em L40S)US$ 0,4081 por execução
Tempo de GPU (L40S)US$ 0,00117 por segundo de GPU
  • Faturado pelo tempo de GPU que a execução realmente leva. Quando a execução começa, 3× o preço típico é reservado do seu saldo e liquidado depois.
  • 1 crédito = US$ 0,01

Calculadora de custos

Calculadora de preços

s

Típico conforme o provedor: cerca de 348,7 s

Total

US$ 40,81

4.081 créditos

Por execução

US$ 0,4081 · 40,81 créditos

Faturado pelo tempo real de GPU; este é uma estimativa.

05

API

Chame CogVideoX-5B (open) com sua chave de API Railwail. Use este ID de modelo na solicitação:

Nenhum exemplo de API verificado

As entradas deste modelo ainda não foram documentadas.

06

Especificações

ID do modelo
cogvideox-5b-open
Desenvolvedor
Zhipu AI
Entrada
Texto
Saída
Vídeo
Faturamento
Por uso (tokens ou tempo de GPU)
Tamanho do modelo
5 billion (also 2B variant)
Licença
Open weights under the CogVideoX Model Licence (free for research and commercial use with attribution).
Entrada do catálogo atualizada
23 de setembro de 2026

Etiquetas

  • zhipu
  • tsinghua
  • cogvideox
  • text-to-video
  • open-weights
  • pricing-tbd
07

Casos de uso

Para que é utilizado

  • Open-source video generation pipelines
  • Academic research on video diffusion
  • On-prem creative tooling
  • LoRA fine-tunes for stylised video
  • Image-to-video animation
  • Benchmark baseline for new video models
08

Perguntas frequentes

O que é CogVideoX-5B (open)?

CogVideoX-5B (open) é um modelo de Zhipu AI na categoria Geração de vídeos.

Quanto custa CogVideoX-5B (open) no Railwail?

No Railwail, CogVideoX-5B (open) custa ≈ US$ 0,4081 por execução. Você é cobrado pelo que cada solicitação realmente usa. O uso é pago com créditos pré-pagos; 1 crédito equivale a US$ 0,01.

Qual é a velocidade de CogVideoX-5B (open)?

Ainda não há execuções medidas suficientes de CogVideoX-5B (open) no Railwail para indicar um tempo de execução. Depende da entrada, das configurações e da carga no provedor.

CogVideoX-5B (open) é melhor que Google Veo 3.1?

Depende da tarefa. CogVideoX-5B (open) (Zhipu AI) e Google Veo 3.1 (Google DeepMind) são ambos modelos na categoria Geração de vídeos. A página de comparação mostra seus preços e especificações lado a lado.

Comparar CogVideoX-5B (open) e Google Veo 3.1
09

Modelos comparáveis

Todos nesta categoria

Todos os modelos através de uma API

Uma chave API para todos os modelos no Railwail. O uso é cobrado a partir de créditos pré-pagos, 1 crédito = US$ 0,01.