Speech-to-Text

Transcribe spoken audio into text.

Available
6of 10
Providers
1
Price range
$0.0012 – $0.156per default run

6 models

  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    whisperstttranscription

    Modalities: Text, Audio

    β‰ˆ $0.0056/run

  • Whisper

    OpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    whisperstttranscription

    Modalities: Text, Audio

    β‰ˆ $0.0034/run

  • SeamlessM4T

    Community

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    metaseamlessstt

    Modalities: Text, Audio

    β‰ˆ $0.156/run

  • Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.

    translationmetaopen-weights

    Modalities: Text, Audio

    β‰ˆ $0.0012/run

  • Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

    whisperstttranscription

    Modalities: Text, Audio

    β‰ˆ $0.0063/run

  • WhisperX

    Community

    WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

    whisperxstttranscription

    Modalities: Text, Audio

    β‰ˆ $0.0241/run

4 models currently unavailable

Their pages stay online, but they can’t be run at the moment.

Speech-to-text models for transcription, meetings, and search

Speech-to-text (STT) models convert spoken audio into written text. The category covers everything from podcast transcripts to real-time captioning pipelines to voice-command interfaces inside mobile apps. Reach for STT when you need to search inside audio, build dictation, summarize meetings, or generate captions for accessibility.

Pricing, trade-offs and pitfalls

Hosted APIs usually bill per minute of audio. On Railwail, the Whisper models run on Replicate and are billed by GPU time; each model page shows the price of a typical run. Some providers charge extra for premium features like speaker diarization, word-level timestamps, summaries, or translation, so do the math with the features you actually need turned on.

The trade-off is accuracy, latency, and feature richness. Whisper Large V3 leads on raw word-error-rate in benchmark evaluations and is open-weights, so you can self-host. Deepgram Nova-3 and AssemblyAI Universal lead on streaming latency (sub-300ms first token) and diarization quality. ElevenLabs Scribe leads on multilingual coverage and code-switching (when speakers swap languages mid-sentence). For batch transcription, Whisper usually wins on cost-and-accuracy. For realtime call transcription, a streaming-first provider wins.

Watch out for noisy audio: word-error-rate roughly doubles below 20 dB SNR on every model, and overlapping speakers degrade diarization even on the flagships. Pre-process with a noise-suppression model (RNNoise, Krisp) if your source is unpredictable. Also watch out for proper nouns: every model still mistranscribes uncommon names, technical terms, and brand names. Most providers accept a `keywords` hint list to bias the decoder β€” use it.

Top picks above cover the most accurate model, the cheapest workhorse, the longest-audio supporter, and the fastest streaming option.

Typical tasks

  • Meeting and call transcription
  • Podcast and video captioning
  • Voice search and dictation
  • Compliance recording analysis
  • Realtime live-captioning
  • Voice-controlled agents

Model comparisons

Frequently asked questions

Which STT model is the most accurate?

Whisper Large V3 leads on word-error-rate in independent benchmarks across most languages. Deepgram Nova-3 leads on English with low-latency streaming. AssemblyAI Universal leads on call-center and meeting audio. Run a sample of your own audio on the model detail page before committing.

Is realtime streaming supported?

Not through Railwail: the transcription endpoint processes a complete audio file. Deepgram, AssemblyAI and OpenAI (Realtime API) offer streaming transcription on their own APIs; for captioning and voice agents, pick a streaming-capable service.

How is STT billed?

Hosted APIs usually bill per minute of audio. On Railwail, the Whisper models are billed by GPU time, with the price of a typical run on each model page. Premium features (diarization, timestamps, translation) sometimes carry surcharges.

What languages are supported?

Whisper Large V3 supports 99 languages. ElevenLabs Scribe covers 100+ with strong code-switching. Deepgram Nova-3 currently covers 40+ with English as the strongest. For lower-resource languages, run a sample first β€” accuracy varies widely.

Can it identify different speakers (diarization)?

Yes on most flagships β€” speaker diarization labels each segment with 'Speaker 1', 'Speaker 2', etc. Accuracy depends on audio quality and how often speakers overlap. Some providers also accept enrollment audio to identify specific named speakers.

Are timestamps provided?

Yes β€” word-level or segment-level timestamps are standard on flagship tiers. Use word-level for video captioning and karaoke-style highlighting; segment-level is enough for transcript search and meeting summaries.

What audio formats are accepted?

MP3, WAV, M4A, FLAC, OGG, and most browser-native streaming formats. Sample rates from 8 kHz (telephony) up to 48 kHz (studio). Max file size varies β€” typically 25 MB on managed APIs and unlimited for self-hosted Whisper.

Can it translate while transcribing?

Yes β€” Whisper has a built-in translate mode that produces English transcripts from any of its 99 supported source languages. ElevenLabs Scribe and a few other providers support translation to a broader target set. Translation accuracy is lower than dedicated translation models β€” fine for search but not for publication.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.