Qwen3 TTS
Alibaba (Qwen)
A unified Text-to-Speech demo featuring three powerful modes: Voice, Clone and Design
text-to-speech
$0.024/1k chars
Turn text into natural-sounding speech, some models with voice cloning.
Quick picks
14 models
Alibaba (Qwen)
A unified Text-to-Speech demo featuring three powerful modes: Voice, Clone and Design
text-to-speech
$0.024/1k chars
Haohe Liu
Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.
audioldmmusic-generationdiffusion
β $0.0157/run
Resemble AI
Resemble AI's open Chatterbox TTS. Zero-shot voice cloning from a short audio prompt with an exaggeration control for emotion intensity, plus CFG weight to balance pacing and fidelity.
ttsvoice-cloningexpressive
$0.030/1k chars
Community
Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.
kokorottsopen-weights
β $0.00030/run
OpenAI
OpenAI's text-to-speech model. Six built-in voices with natural intonation.
fastaffordable
$0.018/1k chars
OpenAI
OpenAI's high-definition TTS model. Better quality for production use cases.
high-quality
$0.036/1k chars
Community
MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.
myshellttsvoice-cloning
β $0.0673/run
Community
Hugging Face Parler-TTS Mini. Lightweight TTS conditioned on a natural-language style description for fine-grained control over voice characteristics.
parlerttsopen-source
β $0.0024/run
Riffusion
Stable-Diffusion-based real-time music generator. Operates on spectrogram images then resynthesizes audio, enables seamless transitions and looping.
music-generationopen-weightsspectrogram
β $0.0576/run
Community
Retrieval-based Voice Conversion. Converts a source recording into a target speaker's voice, preserving pitch, prosody and rhythm.
rvcvoice-conversionvoice-cloning
β $0.0505/run
Community
Spark efficient TTS with disentangled control over speaker, content and style. Strong cross-lingual zero-shot performance.
sparkttsvoice-cloning
β $0.0035/run
Community
Style-based TTS using diffusion and adversarial training. Human-level naturalness in zero-shot voice synthesis from a 3-5s reference clip.
stylettsttsvoice-cloning
β $0.00040/run
Suno
Suno's text-prompted generative audio model. Speech, music, ambient sound and effects with non-verbal cues like laughter or sighs.
barkmusic-generationtts
β $0.0972/run
Community
Multi-voice expressive TTS. Slow but high-quality with strong prosody and natural intonation. Trained for long-form narration use cases.
tortoisettsexpressive
β $0.0816/run
ElevenLabs
ElevenLabs' most natural-sounding TTS model. Supports 29 languages with emotional range.
naturalmultilingual
Currently not offered
Meta
Meta's AudioCraft framework wrapping MusicGen, AudioGen and EnCodec. Unified text-to-audio research toolkit for music and sound effects.
music-generationsound-effectsopen-weights
Deactivated
Other
Cartesia's ultra-low-latency TTS (~90ms TTFB). State-space model with voice cloning support.
cartesiattslow-latency
Currently not offered
Other
Microsoft Edge neural voices accessed via the open-source edge-tts wrapper. 400+ voices across 100+ locales, suitable for batch generation.
microsoftttsmultilingual
Currently not offered
ElevenLabs
ElevenLabs' v3 alpha TTS. Most expressive voice model with audio tags and laughter, higher latency.
ttsexpressivealpha
Deactivated
X-LANCE (SJTU)
Open-source flow-matching TTS with strong zero-shot voice cloning. Code MIT, weights CC-BY-NC.
f5ttsopen-weights
Deactivated
Meta
Meta MusicGen Medium (1.5B params). Strong quality-to-speed tradeoff for text-to-music with optional melody guidance.
music-generationopen-weights
Deactivated
Meta
Meta MusicGen Small (300M params). Fast text-to-music generation suitable for prototyping and low-latency demos.
music-generationopen-weightsfast
Deactivated
Community
MyShell OpenVoice v1. Cross-lingual voice cloning with flexible style control: emotion, accent, rhythm, pauses, and intonation.
myshellttsvoice-cloning
Deactivated
Community
Parler-TTS Large v1. 2.2B parameters, natural-language style prompting and improved prosody over the Mini variant.
parlerttsopen-source
Deactivated
Other
PlayHT's 2.0 generative voice model. Multi-lingual expressive speech synthesis with sub-second latency and high-fidelity voice cloning.
playhtttsvoice-cloning
Currently not offered
Udio
Stability AI's Stable Audio 2.0. Text-to-music up to 3 minutes of full-length, structured tracks at 44.1 kHz.
stabilitymusic-generation
Currently not offered
Community
Coqui's XTTS v2 multilingual TTS with voice cloning from 6 seconds of reference audio. Supports 17 languages and emotion transfer.
coquittsvoice-cloning
Deactivated
Their pages stay online, but they canβt be run at the moment.
Text-to-speech (TTS) models turn written text into natural-sounding spoken audio. The category covers everything from a flat IVR voiceover to expressive narration for audiobooks to realtime conversational agents that hold a phone call. Reach for a TTS model when you need to give software a voice β for accessibility, content production at scale, or conversational AI.
Pricing is usually per character or per thousand characters. On Railwail, OpenAI TTS-1 costs $0.018 and TTS-1 HD $0.036 per 1,000 characters, Chatterbox $0.03; open-weights voices such as Kokoro or OpenVoice v2 are billed by GPU time. A typical short audiobook chapter (3,000 words, around 18,000 characters) therefore costs about $0.32 with TTS-1 or $0.65 with TTS-1 HD. Some providers also charge for voice cloning β a one-time fee to set up a custom voice plus the standard per-character rate at synthesis time.
The trade-off triangle is naturalness, latency, and cost. Flagship voices are nearly indistinguishable from human narration but typically have first-byte latency of 200-600ms, which is fine for batch synthesis but feels sluggish in realtime chat. Streaming TTS (Cartesia, OpenAI Realtime, ElevenLabs Turbo) keeps first-byte latency under 100ms by emitting audio as soon as the first phoneme is decoded. Budget tiers run at flagship speed but with audible robotic artifacts in long sentences.
Watch out for prosody control: even the best models occasionally mis-stress a proper noun, mispronounce an acronym, or lose emotional intent on long sentences. Use SSML tags (where supported) or break long passages into shorter chunks with explicit phrase boundaries. For multilingual content, verify pronunciation on every language pair before shipping β some voices speak English flawlessly and German with a heavy accent.
Top picks above cover the most natural-sounding voice, the cheapest workhorse, the longest-input-supporting model, and the fastest streaming option.
ElevenLabs V3 and Cartesia Sonic currently lead blind A/B tests on naturalness, with OpenAI TTS-HD close behind. The gap narrows for short utterances β under 30 seconds, even budget tiers sound very close to human. Long-form narration is where flagships pull ahead.
On Railwail, OpenAI TTS-1 costs $0.018 per 1,000 characters; open-weights models such as Kokoro, StyleTTS 2 or OpenVoice v2 are billed by GPU time, with the price of a typical run on each model page. Sort the model grid by price for the live ranking.
Yes β most flagship platforms accept a 30-second to 3-minute reference clip and produce a custom voice. Cloning fees vary by provider; synthesis then runs at the standard rate.
Not through Railwail at the moment: the speech endpoint returns the finished audio file. Providers such as Cartesia, ElevenLabs and OpenAI (Realtime API) offer streaming on their own APIs; for interactive agents, use a streaming-capable service.
Flagship platforms cover 30-100 languages with native voices. ElevenLabs V3 ships in 70+, OpenAI TTS in around 50. Quality varies β English, Spanish, German, French, and Mandarin are universally excellent; lower-resource languages can sound robotic or carry accent artifacts.
Modern flagships infer emotion from punctuation and context automatically. For explicit control, use SSML tags (where supported) for emphasis, pauses, and speed; some platforms accept emotion tags like 'excited' or 'calm' directly in the prompt.
MP3 and WAV are universal. PCM, Opus, and Β΅-law are common for telephony. Sample rates run from 16 kHz (telephony) up to 48 kHz (studio). Pick the format that matches your delivery channel.
Almost always yes on commercial tiers β TTS output is treated like a paid voiceover. Cloned voices carry stricter terms: you typically must own or license the source voice. Read the model card for per-provider terms before deploying in ads or paid content.
Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.