Speech-to-Text

Transcribe and understand audio with AI

Modèles reconnaissance vocale pour transcription, réunions et recherche

Les modèles de reconnaissance vocale (STT) convertissent l'audio parlé en texte écrit. La catégorie couvre tout, des transcriptions de podcast aux pipelines de sous-titrage temps réel en passant par les interfaces à commande vocale dans les applis mobiles. On y a recours quand on veut chercher à l'intérieur de l'audio, construire de la dictée, résumer des réunions ou générer des sous-titres pour l'accessibilité.

9 models available

Incredibly Fast Whisper

STTCommunity
Popular

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

€1.00
replicatewhisperstt

Whisper

STTOpenAI
Popular

OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

€2.00
replicateopenaiwhisper

Whisper Large V3

STTOpenAI
Popular

OpenAI's Whisper model. State-of-the-art speech recognition supporting 99+ languages.

€0.305.0s
multilingualpopular

Whisper Large v3 Turbo

STTOpenAI
Popular

OpenAI's distilled Whisper Large v3. ~216x realtime, 99+ languages, MIT-licensed weights.

€0.006
openaiwhisperstt

Deepgram Nova-3

STTCustom

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

€0.004
deepgramstttranscription

SeamlessM4T

STTCommunity

Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

€3.00
replicatemetaseamless

SeamlessM4T v2 Large (Speech)

STTCommunity

Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.

€0.01
replicatetranslationmeta

Whisper Diarization

STTCommunity

Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

€2.00
replicatewhisperstt

WhisperX

STTReplicate

WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

€2.00
replicatewhisperxstt

Top speech-to-text picks

Hand-picked across four common criteria — resolved against the live catalog so the picks track price and performance changes.

Meilleur global
Incredibly Fast Whisper

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

Learn more
Le moins cher
Deepgram Nova-3

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

Learn more
Audio le plus long
Incredibly Fast Whisper

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

Learn more
Le plus rapide
Whisper Large V3

OpenAI's Whisper model. State-of-the-art speech recognition supporting 99+ languages.

Learn more

La tarification est presque toujours à la minute d'audio. Les modèles phares (Whisper Large V3, Deepgram Nova-3, ElevenLabs Scribe) coûtent environ 0,005 à 0,015 € par minute. La transcription d'un podcast d'une heure coûte 0,30 à 0,90 € selon le tiers. Certains fournisseurs facturent en plus des fonctionnalités premium comme la diarisation, les timestamps mot à mot, les résumés ou la traduction, alors faites les calculs avec les fonctionnalités que vous activez vraiment.

Le compromis est précision, latence et richesse fonctionnelle. Whisper Large V3 mène sur le taux d'erreur mot brut dans les évaluations benchmark et est open-weights, donc vous pouvez l'auto-héberger. Deepgram Nova-3 et AssemblyAI Universal mènent sur la latence streaming (premier token sous 300 ms) et la qualité de diarisation. ElevenLabs Scribe mène sur la couverture multilingue et le code-switching (quand les locuteurs alternent les langues en cours de phrase). Pour la transcription par lots, Whisper gagne généralement sur coût-et-précision. Pour la transcription d'appels temps réel, un fournisseur streaming-first gagne.

Attention à l'audio bruité : le taux d'erreur mot environ double sous 20 dB SNR sur chaque modèle, et les locuteurs qui se chevauchent dégradent la diarisation même sur les phares. Pré-traitez avec un modèle de suppression de bruit (RNNoise, Krisp) si votre source est imprévisible. Attention aussi aux noms propres : chaque modèle transcrit encore mal les noms peu communs, les termes techniques et les marques. La plupart des fournisseurs acceptent une liste d'indices `keywords` pour biaiser le décodeur — utilisez-la.

Les top picks ci-dessus couvrent le modèle le plus précis, le cheval de trait le moins cher, le support d'audio le plus long et l'option streaming la plus rapide.

Related comparisons

Side-by-side reviews of the most-compared models in this category.

Frequently asked questions

Start Building with AI

Access all models through a single API. Get free credits when you sign up — no credit card required.