Concept
aka TTS
Text-to-Speech
Models that synthesise spoken audio from written text — the inverse of STT.
Definition
Text-to-speech models — ElevenLabs, Cartesia Sonic, OpenAI's voices — generate audio from text, often with controllable speaker identity, prosody and emotion. Modern systems run end-to-end neural pipelines reaching near-human naturalness.
Common use cases
- Narration
- Voice agents
- Dubbing