Concept
aka TTS

Text-to-Speech

Models that synthesise spoken audio from written text — the inverse of STT.

Definition

Text-to-speech models — ElevenLabs, Cartesia Sonic, OpenAI's voices — generate audio from text, often with controllable speaker identity, prosody and emotion. Modern systems run end-to-end neural pipelines reaching near-human naturalness.

Common use cases

  • Narration
  • Voice agents
  • Dubbing

Related terms

    Text-to-Speech — AI Glossary | Railwail