Speech-to-Text & Transcription Models

Transcription models convert speech into text, many with speaker labels and word-level timestamps.

Use them for subtitles, meeting notes, podcast transcripts and voice interfaces, in dozens of languages.

9 models for this use case

9 models available

Incredibly Fast Whisper

STTReplicate
Popular

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

€1.00
replicatewhisperstt

Whisper

STTReplicate
Popular

OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

€2.00
replicateopenaiwhisper

Whisper Large V3

STTOpenAI
Popular

OpenAI's Whisper model. State-of-the-art speech recognition supporting 99+ languages.

€0.305.0s
multilingualpopular

Whisper Large v3 Turbo

STTOpenAI
Popular

OpenAI's distilled Whisper Large v3. ~216x realtime, 99+ languages, MIT-licensed weights.

€0.006
openaiwhisperstt

Deepgram Nova-3

STTCustom

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

€0.004
deepgramstttranscription

SeamlessM4T

STTReplicate

Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

€3.00
replicatemetaseamless

SeamlessM4T v2 Large (Speech)

STTReplicate

Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.

€0.01
replicatetranslationmeta

Whisper Diarization

STTReplicate

Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

€2.00
replicatewhisperstt

WhisperX

STTReplicate

WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

€2.00
replicatewhisperxstt

Frequently asked questions

One API, pay only for what you use

Try any model with a free generation, no signup. Then 50 free credits and transparent per-use pricing, no subscription.

Related use cases