Google Veo 3.1
Google DeepMind
Latest Veo with image-to-video and context-aware audio
audioi2v
$0.48/s
Generate video clips from text or images and bring stills to life.
Quick picks
46 models
Google DeepMind
Latest Veo with image-to-video and context-aware audio
audioi2v
$0.48/s
Kuaishou (Kling)
Cinematic video up to 15s with multi-shot and native audio
audioi2v
$0.2688/s
Kuaishou (Kling)
Most versatile: multi-reference images, video editing, native audio
audioi2vediting
$0.2688/s
Runway
Top-ranked for motion quality and visual fidelity
top-quality
$0.144/s
Google DeepMind
Google's state-of-the-art video generation model. Simulates real-world physics with various visual styles.
high-quality
$0.60/s
Google DeepMind
Google's Veo 3 served via Replicate. Text-to-video with native synchronized audio generation. High-fidelity motion and scene coherence in short clips.
veotext-to-videoaudio
$0.48/s
Tencent
Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.
hunyuanvideotext-to-video
β $3.06/run
Kuaishou (Kling)
Kuaishou's Kling v2.1, generating 5 and 10 second videos at 720p or 1080p from text or an image. Known for cinematic camera work and realistic physical motion, available on Replicate via the official KwaiVGI account.
videotext-to-videoimage-to-video
$0.060/s
Kuaishou (Kling)
Kuaishou's premium Kling v2.1 Master. Generates 1080p 5s and 10s clips from text or an image with strong dynamics and prompt adherence. The top tier of the Kling 2.1 family.
text-to-videoimage-to-video
$0.336/s
MiniMax
MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.
hailuotext-to-videoimage-to-video
$0.324/video
Runway
Runway's Gen-4 Turbo on Replicate. Fast image-to-video generation producing 5s and 10s clips at 720p with strong character and scene consistency across shots.
gen-4image-to-videofast
$0.060/s
Google DeepMind
Faster cheaper Veo 3 with audio
fastaudio
$0.18/s
Google DeepMind
Faster Veo 3.1 with image-to-video and audio
fastaudioi2v
$0.18/s
xAI video with native audio and lip-sync, up to 15s
audioi2v
$0.060/s
MiniMax
Minimax model for realistic human motion and VFX
i2v1080p
$0.336/video
Kuaishou (Kling)
Kling 2.5 Turbo Pro: Unlock pro-level text-to-video and image-to-video creation with smooth motion, cinematic depth, and remarkable prompt adherence.
kwaivgitext-to-video
$0.084/s
Luma AI
Fast affordable video with I2V support
fastbudgeti2v
$0.072/s
PixVerse
Physics-accurate video generation up to 1080p
i2v1080pphysics
$0.084/s
ByteDance
A faster and cheaper version of Seedance 1 Pro
text-to-video
$0.072/s
ByteDance
Budget ByteDance video, fast and cheap
budgeti2vfast
$0.0432/s
ByteDance
ByteDance video with T2V and I2V, up to 1080p
i2v1080p
$0.18/s
Alibaba (Wan)
The fastest Wan 2.2 text-to-image and image-to-video model
wan-videotext-to-video
$0.030/video
Alibaba (Wan)
Ultra-cheap I2V. Upload image and animate it.
budgeti2vfast
$0.060/video
Alibaba (Wan)
Ultra-cheap T2V for pennies
budgetfast
$0.060/video
Alibaba (Qwen)
(50% off until Aug 30!) - Wan 3.0 generates video from a text prompt or a starting image, with cinematic motion and support for 480p, 720p, and 1080p output up to 30 seconds.
text-to-video
$0.12/s
Community
Plug-and-play motion module that animates personalized Stable Diffusion models without further training. 16-frame clips at 512x512.
animationanimatediffopen-source
β $0.1201/run
Community
ByteDance distillation of AnimateDiff. 4-step sampling for over 10x faster inference at comparable quality to multi-step base model.
animationbytedancefast
$0.00117/GPU s
Community
CogVideoX-5B from Tsinghua/Zhipu AI, an open 5B-parameter text-to-video diffusion transformer. Generates 6-second 720p clips with coherent motion and is widely used in research for its open weights and reproducibility.
cogvideoxzhipuvideo
β $0.60/run
Zhipu AI
Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.
tsinghuacogvideoxtext-to-video
β $0.4081/run
Community
Ant Group EchoMimic. Lifelike audio-driven portrait animation with editable landmark conditioning for fine-grained motion control.
lipsyncant-groupportrait-animation
β $0.48/run
Google Research
Google FILM frame interpolation. Synthesizes high-quality intermediate frames between near-duplicate inputs, designed for large motion gaps.
upscaleframe-interpolationopen-source
β $0.0016/run
Community
Kuaishou LivePortrait. Efficient portrait animation driven by reference videos with stitching, retargeting and motion-control parameters.
lipsynckuaishouportrait-animation
β $0.0949/run
Lightricks
Lightricks' 2B DiT video model. Realtime generation on consumer GPUs (~6s @ H100, 24fps).
ltxtext-to-videoopen-weights
β $0.0229/run
Luma AI
Luma Labs' Ray-2 at 720p on Replicate. Text and image-to-video producing 5s and 9s clips with fast, coherent motion and strong camera control. Successor to Dream Machine.
ray-2text-to-videoimage-to-video
$0.216/s
Community
ByteDance MagicAnimate. Temporally consistent human-image animation driven by a DensePose motion sequence with strong identity preservation.
animationhuman-motionbytedance
β $0.4081/run
MiniMax
MiniMax's video generation model. Fast, high-quality video output with text-to-video capabilities.
fastaffordable
$0.60/video
Genmo
Genmo's Mochi 1, an open text-to-video model with high-fidelity motion built on a 10B Asymmetric Diffusion Transformer. Released under Apache 2.0, it was the largest open video model at launch and is strong on smooth, physically plausible movement.
mochivideotext-to-video
β $0.5041/run
Community
Genmo's 10B open-weights text-to-video model. AsymmDiT architecture, 5.4s @ 480p.
genmomochitext-to-video
β $0.5041/run
Community
Tencent MuseTalk real-time lip-sync model. Audio-driven mouth-region editing in latent space at 30+ fps on a single GPU.
lipsynctencentrealtime
β $0.0624/run
Community
Real-Time Intermediate Flow Estimation. Doubles or quadruples FPS of an existing video via learned optical-flow-based frame interpolation.
upscaleframe-interpolationopen-source
β $0.0444/run
Community
Stylized audio-driven talking-head generator. Synthesizes 3D motion coefficients from audio to animate a single portrait image with natural head movements.
lipsynctalking-headopen-source
β $0.1165/run
Community
SwinIR transformer-based super-resolution and denoising applied per-frame to video. Handles classic, real-world and lightweight upscaling.
upscaletransformeropen-source
β $0.0277/run
Community
Tencent ToonCrafter generative cartoon interpolation model. Synthesizes smooth in-between frames between two cartoon keyframes.
animationtooncrafterinterpolation
β $0.0852/run
Community
Tencent V-Express. Audio-driven portrait animation with progressive training, weak-condition learning, and expressive lip sync.
lipsynctencentportrait-animation
$0.00168/GPU s
Community
Tencent VideoCrafter latent video diffusion. Text-to-video and image-to-video generation up to 2s at 1024x576 with strong motion fidelity.
upscalevideo-generationtencent
β $0.1561/run
Community
Lip-sync model that re-syncs a target video's lip movement to an arbitrary audio track. Robust to identity and language with a lip-sync discriminator loss.
lipsyncvideo-editopen-source
β $0.0067/run
Runway
Runway's latest video generation model. Cinematic quality with precise camera and motion control.
cinematichigh-quality
Deactivated
OpenAI
OpenAI's second-generation Sora video model. Realistic motion, improved physics, audio support.
soratext-to-videoaudio
Deactivated
Image-to-video with synchronized audio using xAI's Grok Imagine Video 1.5 preview model
image-to-video
Deactivated
Alibaba (Qwen)
Alibaba's Happy Horse 1.0 generates videos from text prompts or animates a single image into video. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.
text-to-video
Deactivated
Alibaba (Qwen)
Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images. Supports 720p and 1080p, 3-15 second durations, and five aspect ratios.
text-to-video
Deactivated
Kuaishou (Kling)
Generate 5s and 10s videos in 720p resolution at 30fps
kwaivgiimage-to-video
Deactivated
Kuaishou (Kling)
Generate 5s and 10s videos in 720p resolution at 30fps
kwaivgitext-to-video
Deactivated
Kuaishou (Kling)
Generate 5s and 10s videos in 720p resolution
kwaivgitext-to-video
Deactivated
Kuaishou (Kling)
Kling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation
kwaivgitext-to-video
Deactivated
Lightricks
Delivers high visual fidelity with fast turnaround. Great for daily content creation, marketing teams, and iterative creative workflows.
text-to-video
Deactivated
Luma AI
Luma AI's video generation model. Fast, high-quality video creation with strong physics simulation.
fastphysics
Deactivated
Community
Create 5s 480p videos from a text prompt
leonardoaitext-to-video
Deactivated
Community
Ovi: generate videos with audio from image and text inputs
character-aiimage-to-video
Deactivated
PixVerse
PixVerse's flagship video generation model. Generate cinematic videos with synchronized audio, multi-shot sequences, and precise camera control.
text-to-video
Deactivated
Community
High-fidelity video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
vidutext-to-video
Deactivated
Community
Fast video generation with text-to-video, image-to-video, and start-end-to-video modes. Up to 16 seconds at 1080p with synchronized audio.
vidutext-to-video
Deactivated
Luma AI
Generate 5s and 9s 540p videos
text-to-video
Deactivated
Luma AI
Generate 5s and 9s 540p videos, faster and cheaper than Ray 2
text-to-video
Deactivated
MiniMax
Generate videos with specific camera movements
text-to-video
Deactivated
Alibaba (Wan)
Generate 5s 480p videos. Wan is an advanced and powerful visual generation model developed by Tongyi Lab of Alibaba Group
wan-videotext-to-video
Deactivated
Community
Image-to-video variant of Alibaba's Wan 2.1 14B at 720p, accelerated by WaveSpeedAI. Animates a still input image into a short clip driven by a text prompt, keeping the source composition while adding motion.
wanalibabavideo
Currently not offered
Alibaba (Wan)
Image-to-video at 720p and 480p with Wan 2.2 A14B
wan-videoimage-to-video
Deactivated
Alibaba (Wan)
Alibaba Wan 2.5 Image to video generation with background audio
wan-videoimage-to-video
Deactivated
Alibaba (Wan)
Wan 2.5 image-to-video, optimized for speed
wan-videoimage-to-video
Deactivated
Alibaba (Wan)
Alibaba Wan 2.5 text to video generation model
wan-videotext-to-video
Deactivated
Alibaba (Wan)
Wan 2.5 text-to-video, optimized for speed
wan-videotext-to-video
Deactivated
Alibaba (Wan)
Alibaba Wan 2.6 image to video generation model
wan-videoimage-to-video
Deactivated
Alibaba (Wan)
Alibaba Wan 2.6 text to video generation model
wan-videotext-to-video
Deactivated
Alibaba (Wan)
Generate videos with audio from text prompts using Alibaba's Wan 2.7 model. 1080p, up to 15 seconds, with audio synchronization.
wan-videotext-to-video
Deactivated
Alibaba (Qwen)
Generate videos from text prompts using Alibaba's Wan 3.0 Prime model. Up to 1080p and 30 seconds, with 480p, 720p, and 1080p output.
text-to-video
Deactivated
Alibaba (Wan)
Image-to-video generation with optional audio, multi-shot narrative support, and faster inference
wan-videoimage-to-video
Deactivated
Community
Community fork of AnimateDiff with improved motion modules, beta scheduler control and ControlNet integration for richer animation control.
animationanimatediffopen-source
Deactivated
Community
Champ controllable human image animation. Uses 3D parametric guidance (SMPL) for realistic full-body motion transfer from a single reference image.
animationhuman-motionsmpl
Currently not offered
Community
4D Gaussian-splatting generator extending DreamGaussian to video. Image-conditioned dynamic 3D scenes with view-consistent motion.
animation4dgaussian-splatting
Deactivated
Community
Tencent DynamiCrafter. Animates still images into short videos preserving texture and structure, with strong open-domain coverage.
animationimage-to-videotencent
Currently not offered
Community
Kuaishou's video generation model. Professional quality with strong motion coherence.
professionalmotion
Deactivated
Kuaishou (Kling)
Kuaishou's Kling 1.6 Pro. Premium cinematic motion and physics realism, ~$0.07/sec.
text-to-videoimage-to-videocinematic
Deactivated
Kuaishou (Kling)
Kuaishou's Kling v1.6 Pro on Replicate. Generates 5s and 10s clips in 1080p from text or an image, with cinematic motion and physics realism. The widely used pro tier of the 1.6 generation.
text-to-videoimage-to-videocinematic
Deactivated
Luma's Dream Machine 1.6. 720p text/image-to-video with strong motion and camera control.
lumatext-to-videoimage-to-video
Currently not offered
Community
Motion-Field-Adapter video generator. Controllable image animation from trajectories, keypoints or audio with a strong identity preservation prior.
lipsyncanimationcontrollable
Deactivated
Pika Labs' 2.0 release. Cinematic text/image-to-video with scene composition controls.
text-to-videoimage-to-video
Currently not offered
Community
Real-CUGAN anime-focused upscaler. 2x/3x/4x super-resolution tuned for animation, line-art, and illustrated content.
upscaleanimeopen-source
Deactivated
Runway's faster, cheaper Gen-3 variant. Image-to-video at 5 credits/sec (~$0.05/sec).
runwayimage-to-videofast
Currently not offered
Community
Picsart StreamingT2V. Generates long, consistent videos by chaining short autoregressive clips with motion and appearance memory.
animationlong-formopen-source
Currently not offered
Community
Accelerated inference for Alibaba's Wan 2.1 14B text-to-video at 720p, hosted by WaveSpeedAI on Replicate. Open suite of video foundation models with high-resolution output and faster generation.
alibabawantext-to-video
Currently not offered
Their pages stay online, but they canβt be run at the moment.
Video models turn a prompt β or a still frame, or a short reference clip β into a moving picture. The category is the youngest and most volatile in the catalog: every quarter brings a new flagship that resets the quality bar. Reach for one when you need motion content faster than a human editor can produce it.
Pricing in video is mostly per second of output rather than per token. On Railwail a five-second clip currently costs from $0.30 (Kling v2.1 standard, Runway Gen-4 Turbo) to $3.00 (Veo 2); Veo 3 with audio costs $0.48 per second. Sound-on tiers cost more than silent tiers. Resolution multipliers stack on top of duration: 720p is the standard, 1080p costs roughly 2Γ more, and 4K is rare and expensive.
The trade-off here is duration versus coherence. Most commercial models cap output at five to ten seconds because longer clips drift β characters change clothes, backgrounds morph, and physics breaks down. For longer narratives, generate a sequence of shorter shots and stitch them in post. Image-to-video (start frame + motion prompt) typically produces more stable results than pure text-to-video, especially for characters and product shots.
Watch out for clip length: most models generate between four and ten seconds per call (Kling v2.1: 5 or 10 s, Veo 3.1 Fast: 4, 6 or 8 s), and quality tends to drop at the upper end. If your script needs more, plan for several shots. Also watch out for sound: many models ship silent and you have to overlay audio separately β Veo 3 and Veo 3.1 are examples that can generate audio.
Top picks above cover the flagship realism leader, the cheapest workhorse, the longest-clip model, and the fastest preview option in the category.
Veo 3 and Veo 3.1 are strong on photoreal motion, physics, and integrated audio. Runway Gen-4 and Kling v2.1 are close on visual quality but ship silent. For artistic and stylized output, Pika and Dream Machine often beat flagships at a fraction of the cost.
Most models generate 4 to 10 seconds per call β for example Kling v2.1 5 or 10 seconds, Veo 3.1 Fast 4, 6 or 8 seconds. Longer clips cost more. Beyond 10 seconds you should generate a sequence of shots and edit them together β quality drift dominates over single-call duration today.
Mostly per second of output. On Railwail a 5-second clip currently costs from $0.30 (Kling v2.1 standard, Runway Gen-4 Turbo) to $3.00 (Veo 2); some models such as Hailuo or Wan are priced per clip. Sound-on tiers and 1080p+ resolutions cost more. Every model page shows the price of a standard run.
Yes β image-to-video is the most reliable workflow today. Provide a still frame plus a motion prompt and you get much more stable output than from text alone, especially for character animation and product shots. Most flagships support both modes.
Veo 3 and Veo 3.1 can generate synced audio (dialog, sound effects, music). Many other models output silent video β you generate audio separately with a TTS or music model and overlay it in post. Check the model page for audio support before integrating.
Standard tiers ship 720p. Pro tiers add 1080p at roughly 2Γ the cost. 4K output is rare and expensive in 2026; for finals at higher resolution, upscale in post with a dedicated video upscaler.
Wall-clock time depends on the model: 30 seconds to 2 minutes for a 5-second clip on flagship infrastructure, 5-15 minutes on open-weights shared GPUs. Plan async UX β show progress and let users come back.
Commercial tiers (Veo, Runway, Kling Pro, Pika) grant perpetual royalty-free commercial use. Some open-weights research models restrict to non-commercial β the license is listed on every model page. Read it before you put output in a paid campaign.
Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.