Hosted APIs usually bill per minute of audio. On Railwail, the Whisper models run on Replicate and are billed by GPU time; each model page shows the price of a typical run. Some providers charge extra for premium features like speaker diarization, word-level timestamps, summaries, or translation, so do the math with the features you actually need turned on.
The trade-off is accuracy, latency, and feature richness. Whisper Large V3 leads on raw word-error-rate in benchmark evaluations and is open-weights, so you can self-host. Deepgram Nova-3 and AssemblyAI Universal lead on streaming latency (sub-300ms first token) and diarization quality. ElevenLabs Scribe leads on multilingual coverage and code-switching (when speakers swap languages mid-sentence). For batch transcription, Whisper usually wins on cost-and-accuracy. For realtime call transcription, a streaming-first provider wins.
Watch out for noisy audio: word-error-rate roughly doubles below 20 dB SNR on every model, and overlapping speakers degrade diarization even on the flagships. Pre-process with a noise-suppression model (RNNoise, Krisp) if your source is unpredictable. Also watch out for proper nouns: every model still mistranscribes uncommon names, technical terms, and brand names. Most providers accept a `keywords` hint list to bias the decoder β use it.
Top picks above cover the most accurate model, the cheapest workhorse, the longest-audio supporter, and the fastest streaming option.