Text-to-Speech

Turn text into natural-sounding speech, some models with voice cloning.

Available
14of 27
Providers
6
Price range
$0.00030 – $0.0972per default run

14 models

  • Qwen3 TTS

    Alibaba (Qwen)

    New

    A unified Text-to-Speech demo featuring three powerful modes: Voice, Clone and Design

    text-to-speech

    Modalities: Text, Audio

    $0.024/1k chars

  • AudioLDM 2

    Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    audioldmmusic-generationdiffusion

    Modalities: Text, Audio

    β‰ˆ $0.0157/run

  • Chatterbox

    Resemble AI

    Resemble AI's open Chatterbox TTS. Zero-shot voice cloning from a short audio prompt with an exaggeration control for emotion intensity, plus CFG weight to balance pacing and fidelity.

    ttsvoice-cloningexpressive

    Modalities: Text, Audio

    $0.030/1k chars

  • Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    kokorottsopen-weights

    Modalities: Text, Audio

    β‰ˆ $0.00030/run

  • OpenAI's text-to-speech model. Six built-in voices with natural intonation.

    fastaffordable

    Modalities: Text, Audio

    $0.018/1k chars

  • OpenAI's high-definition TTS model. Better quality for production use cases.

    high-quality

    Modalities: Text, Audio

    $0.036/1k chars

  • OpenVoice v2

    Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    myshellttsvoice-cloning

    Modalities: Text, Audio

    β‰ˆ $0.0673/run

  • Parler-TTS

    Community

    Hugging Face Parler-TTS Mini. Lightweight TTS conditioned on a natural-language style description for fine-grained control over voice characteristics.

    parlerttsopen-source

    Modalities: Text, Audio

    β‰ˆ $0.0024/run

  • Riffusion

    Riffusion

    Stable-Diffusion-based real-time music generator. Operates on spectrogram images then resynthesizes audio, enables seamless transitions and looping.

    music-generationopen-weightsspectrogram

    Modalities: Text, Audio

    β‰ˆ $0.0576/run

  • Retrieval-based Voice Conversion. Converts a source recording into a target speaker's voice, preserving pitch, prosody and rhythm.

    rvcvoice-conversionvoice-cloning

    Modalities: Text, Audio

    β‰ˆ $0.0505/run

  • Spark TTS

    Community

    Spark efficient TTS with disentangled control over speaker, content and style. Strong cross-lingual zero-shot performance.

    sparkttsvoice-cloning

    Modalities: Text, Audio

    β‰ˆ $0.0035/run

  • StyleTTS 2

    Community

    Style-based TTS using diffusion and adversarial training. Human-level naturalness in zero-shot voice synthesis from a 3-5s reference clip.

    stylettsttsvoice-cloning

    Modalities: Text, Audio

    β‰ˆ $0.00040/run

  • Suno's text-prompted generative audio model. Speech, music, ambient sound and effects with non-verbal cues like laughter or sighs.

    barkmusic-generationtts

    Modalities: Text, Audio

    β‰ˆ $0.0972/run

  • Tortoise TTS

    Community

    Multi-voice expressive TTS. Slow but high-quality with strong prosody and natural intonation. Trained for long-form narration use cases.

    tortoisettsexpressive

    Modalities: Text, Audio

    β‰ˆ $0.0816/run

13 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Text-to-speech models for voice apps, audiobooks, and IVR

Text-to-speech (TTS) models turn written text into natural-sounding spoken audio. The category covers everything from a flat IVR voiceover to expressive narration for audiobooks to realtime conversational agents that hold a phone call. Reach for a TTS model when you need to give software a voice β€” for accessibility, content production at scale, or conversational AI.

Pricing, trade-offs and pitfalls

Pricing is usually per character or per thousand characters. On Railwail, OpenAI TTS-1 costs $0.018 and TTS-1 HD $0.036 per 1,000 characters, Chatterbox $0.03; open-weights voices such as Kokoro or OpenVoice v2 are billed by GPU time. A typical short audiobook chapter (3,000 words, around 18,000 characters) therefore costs about $0.32 with TTS-1 or $0.65 with TTS-1 HD. Some providers also charge for voice cloning β€” a one-time fee to set up a custom voice plus the standard per-character rate at synthesis time.

The trade-off triangle is naturalness, latency, and cost. Flagship voices are nearly indistinguishable from human narration but typically have first-byte latency of 200-600ms, which is fine for batch synthesis but feels sluggish in realtime chat. Streaming TTS (Cartesia, OpenAI Realtime, ElevenLabs Turbo) keeps first-byte latency under 100ms by emitting audio as soon as the first phoneme is decoded. Budget tiers run at flagship speed but with audible robotic artifacts in long sentences.

Watch out for prosody control: even the best models occasionally mis-stress a proper noun, mispronounce an acronym, or lose emotional intent on long sentences. Use SSML tags (where supported) or break long passages into shorter chunks with explicit phrase boundaries. For multilingual content, verify pronunciation on every language pair before shipping β€” some voices speak English flawlessly and German with a heavy accent.

Top picks above cover the most natural-sounding voice, the cheapest workhorse, the longest-input-supporting model, and the fastest streaming option.

Typical tasks

  • Audiobook and podcast narration
  • IVR and phone agents
  • Accessibility (screen readers, captions)
  • Conversational AI agents
  • E-learning and explainer videos
  • Voiceover for ads and trailers

Model comparisons

Frequently asked questions

Which TTS model sounds the most human?

ElevenLabs V3 and Cartesia Sonic currently lead blind A/B tests on naturalness, with OpenAI TTS-HD close behind. The gap narrows for short utterances β€” under 30 seconds, even budget tiers sound very close to human. Long-form narration is where flagships pull ahead.

Which is cheapest?

On Railwail, OpenAI TTS-1 costs $0.018 per 1,000 characters; open-weights models such as Kokoro, StyleTTS 2 or OpenVoice v2 are billed by GPU time, with the price of a typical run on each model page. Sort the model grid by price for the live ranking.

Can I clone a specific voice?

Yes β€” most flagship platforms accept a 30-second to 3-minute reference clip and produce a custom voice. Cloning fees vary by provider; synthesis then runs at the standard rate.

Is streaming supported?

Not through Railwail at the moment: the speech endpoint returns the finished audio file. Providers such as Cartesia, ElevenLabs and OpenAI (Realtime API) offer streaming on their own APIs; for interactive agents, use a streaming-capable service.

What languages are supported?

Flagship platforms cover 30-100 languages with native voices. ElevenLabs V3 ships in 70+, OpenAI TTS in around 50. Quality varies β€” English, Spanish, German, French, and Mandarin are universally excellent; lower-resource languages can sound robotic or carry accent artifacts.

Can I control emotion and emphasis?

Modern flagships infer emotion from punctuation and context automatically. For explicit control, use SSML tags (where supported) for emphasis, pauses, and speed; some platforms accept emotion tags like 'excited' or 'calm' directly in the prompt.

What audio formats are output?

MP3 and WAV are universal. PCM, Opus, and Β΅-law are common for telephony. Sample rates run from 16 kHz (telephony) up to 48 kHz (studio). Pick the format that matches your delivery channel.

Is commercial use allowed?

Almost always yes on commercial tiers β€” TTS output is treated like a paid voiceover. Cloned voices carry stricter terms: you typically must own or license the source voice. Read the model card for per-provider terms before deploying in ads or paid content.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.