Video Generation

Generate video clips from text or images and bring stills to life.

Available
46of 91
Providers
15
Price range
$0.0016 – $4.80per default run

46 models

  • Google Veo 3.1

    Google DeepMind

    New

    Latest Veo with image-to-video and context-aware audio

    audioi2v

    Modalities: Video

    $0.48/s

  • Kling v3

    Kuaishou (Kling)

    New

    Cinematic video up to 15s with multi-shot and native audio

    audioi2v

    Modalities: Video

    $0.2688/s

  • Kling v3 Omni

    Kuaishou (Kling)

    New

    Most versatile: multi-reference images, video editing, native audio

    audioi2vediting

    Modalities: Video

    $0.2688/s

  • New

    Top-ranked for motion quality and visual fidelity

    top-quality

    Modalities: Video

    $0.144/s

  • Google Veo 2

    Google DeepMind

    Google's state-of-the-art video generation model. Simulates real-world physics with various visual styles.

    high-quality

    Modalities: Video

    $0.60/s

  • Google's Veo 3 served via Replicate. Text-to-video with native synchronized audio generation. High-fidelity motion and scene coherence in short clips.

    veotext-to-videoaudio

    Modalities: Text, Image, Video

    $0.48/s

  • Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.

    hunyuanvideotext-to-video

    Modalities: Text, Video

    β‰ˆ $3.06/run

  • Kling v2.1

    Kuaishou (Kling)

    Kuaishou's Kling v2.1, generating 5 and 10 second videos at 720p or 1080p from text or an image. Known for cinematic camera work and realistic physical motion, available on Replicate via the official KwaiVGI account.

    videotext-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.060/s

  • Kling v2.1 Master

    Kuaishou (Kling)

    Kuaishou's premium Kling v2.1 Master. Generates 1080p 5s and 10s clips from text or an image with strong dynamics and prompt adherence. The top tier of the Kling 2.1 family.

    text-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.336/s

  • MiniMax Hailuo 02 on Replicate. Text-to-video and image-to-video producing 6s or 10s clips at 768p standard or 1080p pro. Known for accurate real-world physics and stable motion.

    hailuotext-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.324/video

  • Runway's Gen-4 Turbo on Replicate. Fast image-to-video generation producing 5s and 10s clips at 720p with strong character and scene consistency across shots.

    gen-4image-to-videofast

    Modalities: Text, Image, Video

    $0.060/s

  • Google Veo 3 Fast

    Google DeepMind

    New

    Faster cheaper Veo 3 with audio

    fastaudio

    Modalities: Video

    $0.18/s

  • Google Veo 3.1 Fast

    Google DeepMind

    New

    Faster Veo 3.1 with image-to-video and audio

    fastaudioi2v

    Modalities: Video

    $0.18/s

  • New

    xAI video with native audio and lip-sync, up to 15s

    audioi2v

    Modalities: Video

    $0.060/s

  • Hailuo 2.3

    MiniMax

    New

    Minimax model for realistic human motion and VFX

    i2v1080p

    Modalities: Video

    $0.336/video

  • Kling V2.5 Turbo Pro

    Kuaishou (Kling)

    New

    Kling 2.5 Turbo Pro: Unlock pro-level text-to-video and image-to-video creation with smooth motion, cinematic depth, and remarkable prompt adherence.

    kwaivgitext-to-video

    Modalities: Text, Image, Video

    $0.084/s

  • New

    Fast affordable video with I2V support

    fastbudgeti2v

    Modalities: Video

    $0.072/s

  • New

    Physics-accurate video generation up to 1080p

    i2v1080pphysics

    Modalities: Video

    $0.084/s

  • New

    A faster and cheaper version of Seedance 1 Pro

    text-to-video

    Modalities: Text, Image, Video

    $0.072/s

  • Seedance Lite

    ByteDance

    New

    Budget ByteDance video, fast and cheap

    budgeti2vfast

    Modalities: Video

    $0.0432/s

  • Seedance Pro

    ByteDance

    New

    ByteDance video with T2V and I2V, up to 1080p

    i2v1080p

    Modalities: Video

    $0.18/s

  • Wan 2.2 5B Fast

    Alibaba (Wan)

    New

    The fastest Wan 2.2 text-to-image and image-to-video model

    wan-videotext-to-video

    Modalities: Text, Image, Video

    $0.030/video

  • New

    Ultra-cheap I2V. Upload image and animate it.

    budgeti2vfast

    Modalities: Video

    $0.060/video

  • New

    Ultra-cheap T2V for pennies

    budgetfast

    Modalities: Video

    $0.060/video

  • Wan 3

    Alibaba (Qwen)

    New

    (50% off until Aug 30!) - Wan 3.0 generates video from a text prompt or a starting image, with cinematic motion and support for 480p, 720p, and 1080p output up to 30 seconds.

    text-to-video

    Modalities: Text, Image, Video

    $0.12/s

  • AnimateDiff

    Community

    Plug-and-play motion module that animates personalized Stable Diffusion models without further training. 16-frame clips at 512x512.

    animationanimatediffopen-source

    Modalities: Text, Image, Video

    β‰ˆ $0.1201/run

  • ByteDance distillation of AnimateDiff. 4-step sampling for over 10x faster inference at comparable quality to multi-step base model.

    animationbytedancefast

    Modalities: Text, Image, Video

    $0.00117/GPU s

  • CogVideoX-5B

    Community

    CogVideoX-5B from Tsinghua/Zhipu AI, an open 5B-parameter text-to-video diffusion transformer. Generates 6-second 720p clips with coherent motion and is widely used in research for its open weights and reproducibility.

    cogvideoxzhipuvideo

    Modalities: Text, Video

    β‰ˆ $0.60/run

  • Zhipu/Tsinghua's 5B open text-to-video model. 720x480 @ 8fps, 6s clips, image-to-video variant available.

    tsinghuacogvideoxtext-to-video

    Modalities: Text, Image, Video

    β‰ˆ $0.4081/run

  • EchoMimic

    Community

    Ant Group EchoMimic. Lifelike audio-driven portrait animation with editable landmark conditioning for fine-grained motion control.

    lipsyncant-groupportrait-animation

    Modalities: Image, Video, Audio

    β‰ˆ $0.48/run

  • Google FILM frame interpolation. Synthesizes high-quality intermediate frames between near-duplicate inputs, designed for large motion gaps.

    upscaleframe-interpolationopen-source

    Modalities: Image, Video

    β‰ˆ $0.0016/run

  • LivePortrait

    Community

    Kuaishou LivePortrait. Efficient portrait animation driven by reference videos with stitching, retargeting and motion-control parameters.

    lipsynckuaishouportrait-animation

    Modalities: Image, Video

    β‰ˆ $0.0949/run

  • Lightricks' 2B DiT video model. Realtime generation on consumer GPUs (~6s @ H100, 24fps).

    ltxtext-to-videoopen-weights

    Modalities: Text, Image, Video

    β‰ˆ $0.0229/run

  • Luma Labs' Ray-2 at 720p on Replicate. Text and image-to-video producing 5s and 9s clips with fast, coherent motion and strong camera control. Successor to Dream Machine.

    ray-2text-to-videoimage-to-video

    Modalities: Text, Image, Video

    $0.216/s

  • MagicAnimate

    Community

    ByteDance MagicAnimate. Temporally consistent human-image animation driven by a DensePose motion sequence with strong identity preservation.

    animationhuman-motionbytedance

    Modalities: Image, Video

    β‰ˆ $0.4081/run

  • MiniMax's video generation model. Fast, high-quality video output with text-to-video capabilities.

    fastaffordable

    Modalities: Video

    $0.60/video

  • Mochi 1

    Genmo

    Genmo's Mochi 1, an open text-to-video model with high-fidelity motion built on a 10B Asymmetric Diffusion Transformer. Released under Apache 2.0, it was the largest open video model at launch and is strong on smooth, physically plausible movement.

    mochivideotext-to-video

    Modalities: Text, Video

    β‰ˆ $0.5041/run

  • Mochi 1

    Community

    Genmo's 10B open-weights text-to-video model. AsymmDiT architecture, 5.4s @ 480p.

    genmomochitext-to-video

    Modalities: Text, Video

    β‰ˆ $0.5041/run

  • MuseTalk

    Community

    Tencent MuseTalk real-time lip-sync model. Audio-driven mouth-region editing in latent space at 30+ fps on a single GPU.

    lipsynctencentrealtime

    Modalities: Video, Audio

    β‰ˆ $0.0624/run

  • Real-Time Intermediate Flow Estimation. Doubles or quadruples FPS of an existing video via learned optical-flow-based frame interpolation.

    upscaleframe-interpolationopen-source

    Modalities: Video

    β‰ˆ $0.0444/run

  • SadTalker

    Community

    Stylized audio-driven talking-head generator. Synthesizes 3D motion coefficients from audio to animate a single portrait image with natural head movements.

    lipsynctalking-headopen-source

    Modalities: Image, Video, Audio

    β‰ˆ $0.1165/run

  • SwinIR Video

    Community

    SwinIR transformer-based super-resolution and denoising applied per-frame to video. Handles classic, real-world and lightweight upscaling.

    upscaletransformeropen-source

    Modalities: Video

    β‰ˆ $0.0277/run

  • ToonCrafter

    Community

    Tencent ToonCrafter generative cartoon interpolation model. Synthesizes smooth in-between frames between two cartoon keyframes.

    animationtooncrafterinterpolation

    Modalities: Image, Video

    β‰ˆ $0.0852/run

  • V-Express

    Community

    Tencent V-Express. Audio-driven portrait animation with progressive training, weak-condition learning, and expressive lip sync.

    lipsynctencentportrait-animation

    Modalities: Image, Video, Audio

    $0.00168/GPU s

  • VideoCrafter

    Community

    Tencent VideoCrafter latent video diffusion. Text-to-video and image-to-video generation up to 2s at 1024x576 with strong motion fidelity.

    upscalevideo-generationtencent

    Modalities: Text, Image, Video

    β‰ˆ $0.1561/run

  • Wav2Lip

    Community

    Lip-sync model that re-syncs a target video's lip movement to an arbitrary audio track. Robust to identity and language with a lip-sync discriminator loss.

    lipsyncvideo-editopen-source

    Modalities: Video, Audio

    β‰ˆ $0.0067/run

45 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Video generation models for marketing, motion, and prototyping

Video models turn a prompt β€” or a still frame, or a short reference clip β€” into a moving picture. The category is the youngest and most volatile in the catalog: every quarter brings a new flagship that resets the quality bar. Reach for one when you need motion content faster than a human editor can produce it.

Pricing, trade-offs and pitfalls

Pricing in video is mostly per second of output rather than per token. On Railwail a five-second clip currently costs from $0.30 (Kling v2.1 standard, Runway Gen-4 Turbo) to $3.00 (Veo 2); Veo 3 with audio costs $0.48 per second. Sound-on tiers cost more than silent tiers. Resolution multipliers stack on top of duration: 720p is the standard, 1080p costs roughly 2Γ— more, and 4K is rare and expensive.

The trade-off here is duration versus coherence. Most commercial models cap output at five to ten seconds because longer clips drift β€” characters change clothes, backgrounds morph, and physics breaks down. For longer narratives, generate a sequence of shorter shots and stitch them in post. Image-to-video (start frame + motion prompt) typically produces more stable results than pure text-to-video, especially for characters and product shots.

Watch out for clip length: most models generate between four and ten seconds per call (Kling v2.1: 5 or 10 s, Veo 3.1 Fast: 4, 6 or 8 s), and quality tends to drop at the upper end. If your script needs more, plan for several shots. Also watch out for sound: many models ship silent and you have to overlay audio separately β€” Veo 3 and Veo 3.1 are examples that can generate audio.

Top picks above cover the flagship realism leader, the cheapest workhorse, the longest-clip model, and the fastest preview option in the category.

Typical tasks

  • Short-form social ads
  • Product demos and explainers
  • Music videos and stylized shorts
  • Storyboarding for live-action shoots
  • AI-generated B-roll
  • Character animation prototypes

Model comparisons

Frequently asked questions

Which video model is the most realistic?

Veo 3 and Veo 3.1 are strong on photoreal motion, physics, and integrated audio. Runway Gen-4 and Kling v2.1 are close on visual quality but ship silent. For artistic and stylized output, Pika and Dream Machine often beat flagships at a fraction of the cost.

How long can a generated clip be?

Most models generate 4 to 10 seconds per call β€” for example Kling v2.1 5 or 10 seconds, Veo 3.1 Fast 4, 6 or 8 seconds. Longer clips cost more. Beyond 10 seconds you should generate a sequence of shots and edit them together β€” quality drift dominates over single-call duration today.

Is video pricing per-second or per-call?

Mostly per second of output. On Railwail a 5-second clip currently costs from $0.30 (Kling v2.1 standard, Runway Gen-4 Turbo) to $3.00 (Veo 2); some models such as Hailuo or Wan are priced per clip. Sound-on tiers and 1080p+ resolutions cost more. Every model page shows the price of a standard run.

Can I generate video from a starting image?

Yes β€” image-to-video is the most reliable workflow today. Provide a still frame plus a motion prompt and you get much more stable output than from text alone, especially for character animation and product shots. Most flagships support both modes.

Is audio included?

Veo 3 and Veo 3.1 can generate synced audio (dialog, sound effects, music). Many other models output silent video β€” you generate audio separately with a TTS or music model and overlay it in post. Check the model page for audio support before integrating.

What resolutions are supported?

Standard tiers ship 720p. Pro tiers add 1080p at roughly 2Γ— the cost. 4K output is rare and expensive in 2026; for finals at higher resolution, upscale in post with a dedicated video upscaler.

How fast is video generation?

Wall-clock time depends on the model: 30 seconds to 2 minutes for a 5-second clip on flagship infrastructure, 5-15 minutes on open-weights shared GPUs. Plan async UX β€” show progress and let users come back.

Are commercial usage rights granted?

Commercial tiers (Veo, Runway, Kling Pro, Pika) grant perpetual royalty-free commercial use. Some open-weights research models restrict to non-commercial β€” the license is listed on every model page. Read it before you put output in a paid campaign.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.