Concept
aka STT
aka ASR

Speech-to-Text

Models that convert spoken audio into written text — the inverse of TTS.

Definition

Speech-to-text (a.k.a. ASR) models — Whisper, Conformer, Parakeet — take audio waveforms and emit transcripts. They are essential for accessibility, voice assistants, meeting transcription and indexing audio content. Word error rate is the headline metric.

Common use cases

  • Transcription
  • Voice agents
  • Captioning

Related terms

    Speech-to-Text — AI Glossary | Railwail