Metric & Benchmark

Time to First Audio

TTS-specific latency from text-input to first audio packet — the voice analogue of TTFT.

Definition

TTFA measures how quickly a streaming TTS or voice agent begins playing audio after the user finishes speaking. Sub-300ms TTFA is the threshold for natural-feeling conversation, which is why SSM-based TTS like Cartesia Sonic targets it specifically.

Common use cases

  • Voice agents
  • Phone bots
  • Real-time TTS

Related terms

    Time to First Audio — AI Glossary | Railwail