Speech-to-Text

Transcribe and understand audio with AI

Modelos de voz a texto para transcripción, reuniones y búsqueda

Los modelos de voz a texto (STT) convierten el audio hablado en texto escrito. La categoría cubre desde transcripciones de podcasts hasta pipelines de subtitulado en tiempo real e interfaces de comandos de voz dentro de apps móviles. Recurres a STT cuando necesitas buscar dentro de audio, construir dictado, resumir reuniones o generar subtítulos para accesibilidad.

9 models available

Incredibly Fast Whisper

STTCommunity
Popular

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

€1.00
replicatewhisperstt

Whisper

STTOpenAI
Popular

OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

€2.00
replicateopenaiwhisper

Whisper Large V3

STTOpenAI
Popular

OpenAI's Whisper model. State-of-the-art speech recognition supporting 99+ languages.

€0.305.0s
multilingualpopular

Whisper Large v3 Turbo

STTOpenAI
Popular

OpenAI's distilled Whisper Large v3. ~216x realtime, 99+ languages, MIT-licensed weights.

€0.006
openaiwhisperstt

Deepgram Nova-3

STTCustom

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

€0.004
deepgramstttranscription

SeamlessM4T

STTCommunity

Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

€3.00
replicatemetaseamless

SeamlessM4T v2 Large (Speech)

STTCommunity

Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.

€0.01
replicatetranslationmeta

Whisper Diarization

STTCommunity

Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

€2.00
replicatewhisperstt

WhisperX

STTReplicate

WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

€2.00
replicatewhisperxstt

Top speech-to-text picks

Hand-picked across four common criteria — resolved against the live catalog so the picks track price and performance changes.

Mejor en general
Incredibly Fast Whisper

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

Learn more
Más barato
Deepgram Nova-3

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

Learn more
Audio más largo
Incredibly Fast Whisper

Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

Learn more
Más rápido
Whisper Large V3

OpenAI's Whisper model. State-of-the-art speech recognition supporting 99+ languages.

Learn more

La tarificación es casi siempre por minuto de audio. Los modelos punteros (Whisper Large V3, Deepgram Nova-3, ElevenLabs Scribe) cuestan unos 0,005-0,015 € por minuto. La transcripción de un podcast de una hora cuesta 0,30-0,90 € según el nivel. Algunos proveedores cobran extra por funciones premium como diarización de hablantes, marcas de tiempo a nivel de palabra, resúmenes o traducción, así que haz las cuentas con las funciones que realmente vas a activar.

El compromiso es precisión, latencia y riqueza de funciones. Whisper Large V3 lidera en tasa de error de palabra bruta en las evaluaciones benchmark y es de pesos abiertos, así que puedes auto-alojarlo. Deepgram Nova-3 y AssemblyAI Universal lideran en latencia de streaming (primer token por debajo de 300 ms) y calidad de diarización. ElevenLabs Scribe lidera en cobertura multilingüe y code-switching (cuando los hablantes cambian de idioma a mitad de frase). Para transcripción por lotes, Whisper suele ganar en coste y precisión. Para transcripción de llamadas en tiempo real, gana un proveedor streaming-first.

Cuidado con el audio ruidoso: la tasa de error de palabra aproximadamente se duplica por debajo de 20 dB SNR en todos los modelos, y los hablantes que se solapan degradan la diarización incluso en los punteros. Pre-procesa con un modelo de supresión de ruido (RNNoise, Krisp) si tu fuente es impredecible. Cuidado también con los nombres propios: todos los modelos siguen transcribiendo mal nombres poco comunes, términos técnicos y nombres de marca. La mayoría de proveedores aceptan una lista de pistas `keywords` para sesgar el decodificador — úsala.

Las selecciones principales arriba cubren el modelo más preciso, el caballo de batalla más barato, el de mayor soporte de audio y la opción streaming más rápida.

Related comparisons

Side-by-side reviews of the most-compared models in this category.

Frequently asked questions

Start Building with AI

Access all models through a single API. Get free credits when you sign up — no credit card required.