For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
02
Examples
Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
Input
Prompt
Close-up shot of an elderly sailor wearing a yellow raincoat, seated on the deck of a catamaran, slowly puffing on a pipe. His cat lies quietly beside him with eyes closed, enjoying the calm. The warm glow of the setting sun bathes the scene, with gentle waves lapping against the hull and a few seabirds circling slowly above. The camera slowly pushes in, capturing this peaceful and harmonious moment.
Length: 0:05
Settings
resolution
480p
num_frames
81
frames_per_second
16
Input
Prompt
she looks up
Length: 0:05
Settings
resolution
480p
num_frames
81
frames_per_second
16
Input
Prompt
she looks up, and then looks back
Length: 0:05
Settings
resolution
480p
num_frames
81
frames_per_second
16
Example prompts
Examples from the Railwail catalog. They were not generated live on this page.
Animate
Length: 0:05
Subject turns head and smiles
03
About Wan 2.2 Image-to-Video
TL;DRAs of September 23, 2026
Wan 2.2 Image-to-Video is a model by Alibaba (Wan) in the Video generation category. On Railwail, Wan 2.2 Image-to-Video costs $0.060 per video. Depending on the settings, the price ranges from $0.060/video to $0.174/video.
Background
About Alibaba (Tongyi Wanxiang Lab)
Founded 1999 ยท Hangzhou, China
Alibaba's Tongyi Lab runs the Qwen LLM family and the Wanxiang generative-media family (image and video). After Wan 2.0 (mid-2024) and Wan 2.1 (early 2025), the team released Wan 2.2 in 2025 as the next-generation open-weight video model. Wan 2.2 ships as two purpose-tuned variants -- Image-to-Video (Wan-I2V) and Text-to-Video (Wan-T2V) -- alongside an audio/A2V variant. Wan 2.2 Image-to-Video is positioned for animating a user-supplied first frame with strong identity preservation and rich motion, and remains under the permissive Wan-series licence that enables broad commercial and research use.
Diffusion Transformer (DiT) with image-conditioning adapters and 3D causal VAE
Wan 2.2 Image-to-Video is a Diffusion Transformer operating on a 3D causal Wan-VAE latent. The model is initialised from a Wan 2.2 base and fine-tuned for image-to-video with dedicated conditioning adapters that inject features from the user-provided first frame at multiple resolutions of the DiT, ensuring strong identity preservation and continuity. Text conditioning uses a Qwen-family multilingual encoder with strong Chinese-English capability. Native generation is 5 seconds at 720p / 24 fps (with 1080p in extended modes). Wan 2.2's MoE-style scaling and improved 3D RoPE enable better motion physics and reduced identity drift relative to Wan 2.1. The training recipe includes a curriculum mixing image-to-image, image-to-short-video and image-to-long-video data with synthetic dense captions. The release is open-weight on Hugging Face and GitHub under the Wan-series permissive licence.
Parameters
14 billion (I2V flagship); smaller variants available
Capabilities
Open-weight image-to-video model with strong identity preservation
Bilingual Chinese/English prompts via Qwen-based text encoder
Identity-preserving conditioning on a user first frame
Permissive Wan-series licence for research and commercial use
Active community ecosystem on Hugging Face / ComfyUI
Compatible with LoRA fine-tunes and reference adapters
Reproducible training recipe documented in technical reports
Best for: open-source image animation pipelines, e-commerce product motion, character animation.
Training & license
Curated multi-million-clip multilingual video corpus with paired first-frame conditioning data and dense bilingual captions; specifics documented in the Wan 2.2 technical materials.
License: Open weights under an Apache-style permissive licence (Wan-series release).
Safety testing: Data-level filtering and recommended NSFW classifier for deployment; Alibaba-hosted endpoints follow Chinese CAC algorithmic-content rules.
Known limitations
Native duration 5 seconds
Identity can drift on long extensions
No native audio
High VRAM requirements for the 14B flagship
Closed leaders (Veo 3, Sora 2, Kling v3) still ahead on absolute fidelity
curl https://railwail.com/api/v1/videos/generations \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan-i2v",
"prompt": "A slow drone shot over a misty pine forest at sunrise",
"start_image": "https://example.com/first-frame.jpg"
}'
# Response: {"job_id": "...", "status": "queued", ...}
# Poll until status is completed, failed or cancelled:
curl https://railwail.com/api/v1/jobs/JOB_ID \
-H "Authorization: Bearer $RAILWAIL_API_KEY"
import os
import time
import requests
API = "https://railwail.com/api/v1"
headers = {"Authorization": f"Bearer {os.environ['RAILWAIL_API_KEY']}"}
job = requests.post(
f"{API}/videos/generations",
headers=headers,
json={
"model": "wan-i2v",
"prompt": "A slow drone shot over a misty pine forest at sunrise",
"start_image": "https://example.com/first-frame.jpg",
},
).json()
while True:
status = requests.get(f"{API}/jobs/{job['job_id']}", headers=headers).json()
if status["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(5)
print(status["status"], status.get("output_url"))
const API = "https://railwail.com/api/v1";
const headers = {
Authorization: `Bearer ${process.env.RAILWAIL_API_KEY}`,
"Content-Type": "application/json",
};
const job = await fetch(`${API}/videos/generations`, {
method: "POST",
headers,
body: JSON.stringify({
model: "wan-i2v",
prompt: "A slow drone shot over a misty pine forest at sunrise",
start_image: "https://example.com/first-frame.jpg"
}),
}).then((r) => r.json());
let status;
do {
await new Promise((r) => setTimeout(r, 5000));
status = await fetch(`${API}/jobs/${job.job_id}`, { headers }).then((r) => r.json());
} while (!["completed", "failed", "cancelled"].includes(status.status));
console.log(status.status, status.output_url);
Wan 2.2 Image-to-Video is a model by Alibaba (Wan) in the Video generation category. On Railwail you can call it with an API key through the Railwail API.
How much does Wan 2.2 Image-to-Video cost on Railwail?
On Railwail, Wan 2.2 Image-to-Video costs $0.060 per video. Depending on the settings, the price ranges from $0.060/video to $0.174/video. The price is known before the run starts. Usage is paid from prepaid credits; 1 credit equals $0.01.
Which settings does Wan 2.2 Image-to-Video support?
According to its input schema, Wan 2.2 Image-to-Video knows these parameters: prompt, start_image, and end_image.
How fast is Wan 2.2 Image-to-Video?
There are not enough measured runs of Wan 2.2 Image-to-Video on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Wan 2.2 Image-to-Video better than Google Veo 3.1?
That depends on the task. Wan 2.2 Image-to-Video (Alibaba (Wan)) and Google Veo 3.1 (Google DeepMind) are both models in the Video generation category. The comparison page shows their prices and specifications side by side.
How do I use Wan 2.2 Image-to-Video through the API?
Create a Railwail API key and send your request with the model ID wan-i2v. Code examples for curl, Python and JavaScript are in the API section of this page.
MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.