HunyuanVideo

Video generationAvailable
by TencentModel ID: hunyuan-video

Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.

Price
โ‰ˆ $3.06/run
Input โ†’ output
Text โ†’ Video
Developer
Tencent
Updated
September 23, 2026
01

Playground

Try HunyuanVideo

Input & output

โ‰ˆ $3.06/run
Try HunyuanVideo

0 / 2,000

What video to generate

Advanced settings (1)

Optional seed for reproducibility

Output
Your video appears here.

This run

about $3.06 ยท 306 credits

$9.18 (918 credits) are reserved at the start; the actual GPU time is billed.

For accounts without a purchase: runs above 2 credits need a top-up.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    A cat walks on the grass, realistic style

    Length: 0:05

    Settings

    width
    864
    height
    480
    fps
    24
  • Prompt

    Close-up, A little girl wearing a red hoodie in winter strikes a match. The sky is dark, there is a layer of snow on the ground, and it is still snowing lightly. The flame of the match flickers, illuminating the girl's face intermittently.

    Length: 0:05

    Settings

    width
    864
    height
    480
    fps
    24
  • Prompt

    a stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage. She wears a black leather jacket, a long red dress, and black boots, and carries a black purse. She wears sunglasses and red lipstick. She walks confidently and casually. The street is damp and reflective, creating a mirror effect of the colorful lights. Many pedestrians walk about.

    Length: 0:05

    Settings

    width
    864
    height
    480
    fps
    24

Example prompts

Examples from the Railwail catalog. They were not generated live on this page.

  • Underwater Scene

    Length: 0:08

    Camera gliding through a vibrant coral reef teeming with tropical fish, sunlight filtering through the crystal-clear water creating dancing light patterns on the ocean floor, documentary style
  • Fantasy Animation

    Length: 0:06

    A tiny glowing fairy emerging from an opening flower bud in an enchanted forest, sparkles trailing behind her wings as she takes flight, magical atmosphere with bioluminescent plants
03

About HunyuanVideo

TL;DRAs of September 23, 2026

HunyuanVideo is a model by Tencent in the Video generation category. On Railwail, HunyuanVideo costs โ‰ˆ $3.06 per run.

HunyuanVideo from Tencent is a large open-source text-to-video model built on a DiT architecture with a unified full-attention design. It is known for realistic physics, large-scale motion and good text alignment, and ships its own weights, making it a popular base for fine-tuning and LoRA training.

Background

About Tencent

Founded 1998 ยท Shenzhen, China

Tencent Holdings is one of the largest technology and entertainment conglomerates in the world, founded in 1998 by Ma Huateng (Pony Ma) and four co-founders in Shenzhen. Tencent's AI Lab, founded in 2016, and the Hunyuan team (the company's foundation-model group) developed the Hunyuan family of text, image, 3D and video models. HunyuanVideo (released December 2024 with the 13B foundation model open-sourced) was the largest publicly released open-weight video diffusion model at the time, and rapidly became a popular base for community fine-tunes (LoRAs, control nets, audio-driven extensions). The model is hosted on Hugging Face and GitHub under a custom licence that allows research and limited commercial use.

Visit Tencent

Architecture

Diffusion Transformer with dual-stream design (separate text/video streams)

HunyuanVideo is a 13B-parameter Diffusion Transformer (DiT) operating on a 3D causal VAE latent. Its denoiser uses a dual-stream / single-stream design inspired by FLUX-1: dedicated text and video streams first process their modalities separately and then concatenate for joint self-attention. Text conditioning combines a CLIP-like vision-language encoder and a multilingual large-language-model text encoder (MLLM) for stronger prompt understanding, especially on long captions. The model is trained with Flow Matching and 3D Rotary Position Embeddings, and uses a progressive curriculum from images to short videos to long videos at 720p / 24 fps. The accompanying paper details a meticulously curated multi-billion-clip dataset with hierarchical filtering and dense bilingual captions. HunyuanVideo supports text-to-video natively and the open release includes a high-quality 5-second / 720p model; community follow-ups added image-to-video, audio-driven and longer-duration variants.

Parameters
13 billion

Capabilities

  • 13B open-weight text-to-video model (one of the largest released openly)
  • Generates ~5-second clips at 720p / 24 fps natively
  • Bilingual Chinese/English prompt handling via MLLM text encoder
  • Dual-stream DiT architecture inspired by FLUX-1
  • Strong prompt adherence on long, dense captions
  • Active community ecosystem (image-to-video, LoRAs, audio-sync)
  • Runs on 60-80 GB GPU memory (FP16), 40 GB with optimisations
  • Permissive licence for non-commercial research and limited commercial use
  • Best for: open-source video pipelines, research, on-prem creative tooling.

Training & license

Multi-billion-clip curated video corpus with hierarchical aesthetic, motion and caption-quality filtering, plus dense bilingual captions produced by an in-house MLLM captioner.

License: Tencent Hunyuan Community Licence: weights free for research and limited commercial use with attribution; restrictions apply above thresholds and for certain jurisdictions.

Safety testing: Released with safety filtering during training and a recommended NSFW classifier for deployment; subject to Chinese CAC algorithmic-content rules on Tencent-hosted surfaces.

Known limitations

  • Native duration ~5 seconds
  • 720p only (extensions require upscalers)
  • High VRAM requirements vs smaller open models
  • No native audio
  • Licence has thresholds for very large commercial deployers
04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 418 s on 4x H100)$3.06 per run
GPU time (4x H100)$0.00732 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 418 s

Total

$306.00

30,600 credits

Per run

$3.06 ยท 306 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call HunyuanVideo with your Railwail API key. Use this model ID in the request:
curl https://railwail.com/api/v1/videos/generations \
  -H "Authorization: Bearer $RAILWAIL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "hunyuan-video",
    "prompt": "A slow drone shot over a misty pine forest at sunrise"
  }'

# Response: {"job_id": "...", "status": "queued", ...}
# Poll until status is completed, failed or cancelled:
curl https://railwail.com/api/v1/jobs/JOB_ID \
  -H "Authorization: Bearer $RAILWAIL_API_KEY"
Set your key as RAILWAIL_API_KEYCreate API key

Videos run asynchronously: the first call returns a job_id; query /api/v1/jobs/JOB_ID until the status is completed.

06

Specifications

Model ID
hunyuan-video
Developer
Tencent
Input
Text
Output
Video
Billing
By usage (tokens or GPU time)
Model size
13 billion
License
Tencent Hunyuan Community Licence: weights free for research and limited commercial use with attribution; restrictions apply above thresholds and for certain jurisdictions.
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • promptrequired

    What video to generate

    Type: Text
    Default: โ€“
    Allowed values: up to 2,000 characters
  • seed

    Optional seed for reproducibility

    Type: Integer
    Default: โ€“
    Allowed values: โ€“

Tags

  • replicate
  • tencent
  • hunyuan
  • video
  • text-to-video
  • open-weights
07

Use cases

What it is used for

  • Open-source text-to-video pipelines
  • Research on large video DiTs
  • On-prem creative tools for studios
  • LoRA fine-tunes for branded styles
  • Image-to-video via community extensions
  • Audio-driven talking-head extensions
08

Frequently asked questions

What is HunyuanVideo?

HunyuanVideo is a model by Tencent in the Video generation category. On Railwail you can call it with an API key through the Railwail API.

How much does HunyuanVideo cost on Railwail?

On Railwail, HunyuanVideo costs โ‰ˆ $3.06 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

Which settings does HunyuanVideo support?

According to its input schema, HunyuanVideo knows these parameters: prompt (up to 2,000 characters) and seed.

How fast is HunyuanVideo?

There are not enough measured runs of HunyuanVideo on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is HunyuanVideo better than Google Veo 3.1?

That depends on the task. HunyuanVideo (Tencent) and Google Veo 3.1 (Google DeepMind) are both models in the Video generation category. The comparison page shows their prices and specifications side by side.

Compare HunyuanVideo and Google Veo 3.1

How do I use HunyuanVideo through the API?

Create a Railwail API key and send your request with the model ID hunyuan-video. Code examples for curl, Python and JavaScript are in the API section of this page.

09

Comparable models

All in this category

Use HunyuanVideo via the API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.