WhisperX

Speech-to-textAvailable
by CommunityModel ID: whisperx

WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

Price
β‰ˆ $0.0241/run
Input β†’ output
Audio β†’ Text
Developer
Community
Updated
September 23, 2026
01

Playground

Try WhisperX

Input & output

β‰ˆ $0.0241/run
Try WhisperX

URL or upload of the audio file to transcribe

Optional ISO-639-1 language code; auto-detected if omitted

Advanced settings (3)
Output
The transcript appears here.

This run

about $0.0241 Β· 2.41 credits

$0.0721 (7.21 credits) are reserved at the start; the actual GPU time is billed.

For accounts without a purchase: runs above 2 credits need a top-up.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Output (JSON, shortened)

    {
      "segments": [
        {
          "end": 30.811,
          "text": "The little tales they tell are false. The door was barred, locked and bolted as well. Ripe pears are fit for a queen's table. A big wet stain was on the round carpet. The kite dipped and swayed but stayed aloft. The pleasant hours fly by much too soon. The room was crowded with a mild wob.",
          "start": 2.585
        },
        {
          "end": 48.592,
          "text": "The room was crowded with a wild mob. This strong arm shall shield your honor. She blushed when he gave her a white orchid. The beetle droned in the hot June sun.",
          "start": 33.029
        }
      ],
      "detected_language": "en"
    }
03

About WhisperX

TL;DRAs of September 23, 2026

WhisperX is a model by Community in the Speech-to-text category. On Railwail, WhisperX costs β‰ˆ $0.0241 per run.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (β‰ˆ 14 s on A100 (80GB))$0.0241 per run
GPU time (A100 (80GB))$0.00168 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3Γ— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 14.3 s

Total

$2.41

241 credits

Per run

$0.0241 Β· 2.41 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call WhisperX with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
whisperx
Developer
Community
Input
Audio
Output
Text
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • audio_filerequired

    URL or upload of the audio file to transcribe

    Type: Text
    Default: –
    Allowed values: –
  • language

    Optional ISO-639-1 language code; auto-detected if omitted

    Type: Text
    Default: –
    Allowed values: –
  • batch_size
    Type: Integer
    Default: 64
    Allowed values: 1 to 64
  • diarization
    Type: Yes/no
    Default: false
    Allowed values: –
  • align_output
    Type: Yes/no
    Default: true
    Allowed values: –

Tags

  • replicate
  • whisperx
  • stt
  • transcription
  • diarization
  • word-timestamps
  • multilingual
07

Use cases

08

Frequently asked questions

What is WhisperX?

WhisperX is a model by Community in the Speech-to-text category.

How much does WhisperX cost on Railwail?

On Railwail, WhisperX costs β‰ˆ $0.0241 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

Which settings does WhisperX support?

According to its input schema, WhisperX knows these parameters: audio_file, language, batch_size (1 to 64), diarization, and align_output.

How fast is WhisperX?

There are not enough measured runs of WhisperX on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is WhisperX better than Incredibly Fast Whisper?

That depends on the task. WhisperX (Community) and Incredibly Fast Whisper (Community) are both models in the Speech-to-text category. The comparison page shows their prices and specifications side by side.

Compare WhisperX and Incredibly Fast Whisper
09

Comparable models

All in this category
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    β‰ˆ $0.0056/run

    77 % cheaper per unit

    Compare WhisperX vs. Incredibly Fast Whisper
  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    β‰ˆ $0.0034/run

    86 % cheaper per unit

    Compare WhisperX vs. Whisper
  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    β‰ˆ $0.156/run

    547 % more expensive per unit

    Compare WhisperX vs. SeamlessM4T

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.