WhisperX
whisperxWhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.
- Price
- โ $0.0241/run
- Input โ output
- Audio โ Text
- Developer
- Community
- Updated
- September 23, 2026
Playground
Try WhisperX
Input & output
This run
about $0.0241 ยท 2.41 credits
$0.0721 (7.21 credits) are reserved at the start; the actual GPU time is billed.
For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
Examples
Output (JSON, shortened)
{ "segments": [ { "end": 30.811, "text": "The little tales they tell are false. The door was barred, locked and bolted as well. Ripe pears are fit for a queen's table. A big wet stain was on the round carpet. The kite dipped and swayed but stayed aloft. The pleasant hours fly by much too soon. The room was crowded with a mild wob.", "start": 2.585 }, { "end": 48.592, "text": "The room was crowded with a wild mob. This strong arm shall shield your honor. She blushed when he gave her a white orchid. The beetle droned in the hot June sun.", "start": 33.029 } ], "detected_language": "en" }
About WhisperX
WhisperX is a model by Community in the Speech-to-text category. On Railwail, WhisperX costs โ $0.0241 per run.
Pricing
| Typical run (โ 14 s on A100 (80GB)) | $0.0241 per run |
|---|---|
| GPU time (A100 (80GB)) | $0.00168 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 14.3 s
Total
$2.41
241 credits
Per run
$0.0241 ยท 2.41 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
whisperx- Developer
- Community
- Category
- Speech-to-text
- Input
- Audio
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
audio_filerequiredURL or upload of the audio file to transcribe
Type: TextDefault: โAllowed values: โlanguageOptional ISO-639-1 language code; auto-detected if omitted
Type: TextDefault: โAllowed values: โbatch_sizeType: IntegerDefault:64Allowed values: 1 to 64diarizationType: Yes/noDefault:falseAllowed values: โalign_outputType: Yes/noDefault:trueAllowed values: โ
Tags
- replicate
- whisperx
- stt
- transcription
- diarization
- word-timestamps
- multilingual
Use cases
Frequently asked questions
What is WhisperX?
WhisperX is a model by Community in the Speech-to-text category.
How much does WhisperX cost on Railwail?
On Railwail, WhisperX costs โ $0.0241 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
Which settings does WhisperX support?
According to its input schema, WhisperX knows these parameters: audio_file, language, batch_size (1 to 64), diarization, and align_output.
How fast is WhisperX?
There are not enough measured runs of WhisperX on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is WhisperX better than Incredibly Fast Whisper?
That depends on the task. WhisperX (Community) and Incredibly Fast Whisper (Community) are both models in the Speech-to-text category. The comparison page shows their prices and specifications side by side.
Compare WhisperX and Incredibly Fast WhisperComparable models
All in this category- Incredibly Fast WhisperCommunity
Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.
- WhisperOpenAI
OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.
- SeamlessM4TCommunity
Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.