For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
02
Examples
Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
Prompt
A sports car is driving very fast along a beach at sunset
Length: 0:05
Settings
aspect_ratio
16:9
resolution
480p
num_frames
81
frames_per_second
16
Prompt
A sports car is driving very fast along a beach at sunset
Length: 0:05
Settings
aspect_ratio
16:9
resolution
480p
num_frames
81
frames_per_second
16
Prompt
A sports car is driving very fast along a beach at sunset
Length: 0:05
Settings
aspect_ratio
16:9
resolution
480p
num_frames
81
frames_per_second
16
Example prompts
Examples from the Railwail catalog. They were not generated live on this page.
Quick
Length: 0:05
Cat playing with yarn on wooden floor
03
About Wan 2.2 Text-to-Video
TL;DRAs of September 23, 2026
Wan 2.2 Text-to-Video is a model by Alibaba (Wan) in the Video generation category. On Railwail, Wan 2.2 Text-to-Video costs $0.060 per video. Depending on the settings, the price ranges from $0.060/video to $0.12/video. Supported aspect ratios: 16:9 and 9:16.
Background
About Alibaba (Tongyi Wanxiang Lab)
Founded 1999 ยท Hangzhou, China
Alibaba's Tongyi Lab in Hangzhou runs the Qwen LLM family and the Wanxiang generative-media family. After Wan 2.0 (mid-2024) and Wan 2.1 (early 2025), the team released Wan 2.2 in 2025 as the next-generation open-weight video model. Wan 2.2 ships as purpose-tuned variants for Text-to-Video, Image-to-Video and Audio/A2V. Wan 2.2 Text-to-Video is the flagship pure-text-conditioned variant and replaces Wan 2.1 T2V-14B as the principal open-weight text-to-video reference for the Chinese research community. The Wan team consistently rank near the top of open-model VBench leaderboards and ship reproducible training code under a permissive Wan-series licence.
Diffusion Transformer (DiT) with 3D causal VAE; MoE-style scaling
Wan 2.2 Text-to-Video is a Diffusion Transformer operating on a 3D causal Wan-VAE latent. Wan 2.2 introduces architectural refinements over Wan 2.1: improved 3D Rotary Position Embeddings, larger attention windows, and (in the flagship) a Mixture-of-Experts feed-forward design that routes tokens to specialist experts. Text conditioning uses a Qwen-family multilingual encoder with strong Chinese-English capability. The denoiser is trained with Flow Matching on a curated multi-million-clip multilingual video corpus with synthetic dense bilingual captions. Native generation is 5 seconds at 720p / 24 fps (with 1080p extensions). The training recipe and weights are open-source on Hugging Face and GitHub under the Wan-series permissive licence, designed to enable broad commercial and research use.
Parameters
14 billion (flagship); smaller variants available
Capabilities
Open-weight text-to-video flagship at 14B parameters (smaller variants available)
Bilingual Chinese/English prompts via Qwen-based text encoder
MoE-style scaling and improved 3D RoPE in flagship variant
Permissive Wan-series licence for research and commercial use
Top-tier results on VBench among open-weight models
Active community ecosystem (LoRAs, fine-tunes, ComfyUI nodes)
Reproducible training recipe and code
Best for: open-source video pipelines, research, on-prem creative tooling, branded fine-tunes.
Training & license
Curated multi-million-clip multilingual video corpus filtered for aesthetics, motion and caption quality, with dense bilingual captions; specifics documented in Wan technical materials.
License: Open weights under an Apache-style permissive licence (Wan-series release).
Safety testing: Data-level filtering and recommended NSFW classifier for deployment; Alibaba-hosted endpoints follow Chinese CAC algorithmic-content rules.
Known limitations
Native duration 5 seconds
No native audio
High VRAM requirements for the 14B flagship
Closed leaders (Veo 3, Sora 2, Kling v3) still ahead on absolute fidelity
Resolution capped at 720p natively (1080p only in extended modes)
curl https://railwail.com/api/v1/videos/generations \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-t2v",
"prompt": "A slow drone shot over a misty pine forest at sunrise"
}'
# Response: {"job_id": "...", "status": "queued", ...}
# Poll until status is completed, failed or cancelled:
curl https://railwail.com/api/v1/jobs/JOB_ID \
-H "Authorization: Bearer $RAILWAIL_API_KEY"
import os
import time
import requests
API = "https://railwail.com/api/v1"
headers = {"Authorization": f"Bearer {os.environ['RAILWAIL_API_KEY']}"}
job = requests.post(
f"{API}/videos/generations",
headers=headers,
json={
"model": "wan-t2v",
"prompt": "A slow drone shot over a misty pine forest at sunrise",
},
).json()
while True:
status = requests.get(f"{API}/jobs/{job['job_id']}", headers=headers).json()
if status["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(5)
print(status["status"], status.get("output_url"))
const API = "https://railwail.com/api/v1";
const headers = {
Authorization: `Bearer ${process.env.RAILWAIL_API_KEY}`,
"Content-Type": "application/json",
};
const job = await fetch(`${API}/videos/generations`, {
method: "POST",
headers,
body: JSON.stringify({
model: "wan-t2v",
prompt: "A slow drone shot over a misty pine forest at sunrise"
}),
}).then((r) => r.json());
let status;
do {
await new Promise((r) => setTimeout(r, 5000));
status = await fetch(`${API}/jobs/${job.job_id}`, { headers }).then((r) => r.json());
} while (!["completed", "failed", "cancelled"].includes(status.status));
console.log(status.status, status.output_url);
Wan 2.2 Text-to-Video is a model by Alibaba (Wan) in the Video generation category. On Railwail you can call it with an API key through the Railwail API.
How much does Wan 2.2 Text-to-Video cost on Railwail?
On Railwail, Wan 2.2 Text-to-Video costs $0.060 per video. Depending on the settings, the price ranges from $0.060/video to $0.12/video. The price is known before the run starts. Usage is paid from prepaid credits; 1 credit equals $0.01.
Which settings does Wan 2.2 Text-to-Video support?
According to its input schema, Wan 2.2 Text-to-Video knows these parameters: prompt and aspect_ratio (16:9 or 9:16).
How fast is Wan 2.2 Text-to-Video?
There are not enough measured runs of Wan 2.2 Text-to-Video on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Wan 2.2 Text-to-Video better than Google Veo 3.1?
That depends on the task. Wan 2.2 Text-to-Video (Alibaba (Wan)) and Google Veo 3.1 (Google DeepMind) are both models in the Video generation category. The comparison page shows their prices and specifications side by side.
How do I use Wan 2.2 Text-to-Video through the API?
Create a Railwail API key and send your request with the model ID wan-t2v. Code examples for curl, Python and JavaScript are in the API section of this page.
MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.