Whisper Diarization

Speech-to-textAvailable
by CommunityModel ID: whisper-diarization

Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

Price
β‰ˆ $0.0063/run
Input β†’ output
Audio β†’ Text
Developer
Community
Updated
September 23, 2026
01

Playground

Try Whisper Diarization

Input & output

β‰ˆ $0.0063/run
Try Whisper Diarization

0 / 1,000

Optional vocabulary or context hint

URL of the audio file to transcribe and diarize

Optional ISO-639-1 language code; auto-detected if omitted

Advanced settings (2)

Optional fixed speaker count; estimated if omitted

Output
The transcript appears here.

This run

about $0.0063 Β· 0.63 credits

$0.0188 (1.88 credits) are reserved at the start; the actual GPU time is billed.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 5 runs of this model.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    LLama, AI, Meta.

    Output (JSON, shortened)

    {
      "language": "en",
      "segments": [
        {
          "end": 4.48,
          "text": "Let me ask you about AI.",
          "start": 2.94,
          "words": [
            {
              "end": 3.12,
              "word": "Let",
              "start": 2.94,
              "speaker": "SPEAKER_01",
              "probability": 0.685546875
            },
            {
              "end": 3.26,
              "word": "me",
              "start": 3.12,
              "speaker": "SPEAKER_01",
              "probability": 0.9990234375
            },
            {
              "end": 3.74,
              "word": "ask",
              "start": 3.26,
              "speaker": "SPEAKER_01",
              "probability": 0.998046875
            },
            {
              "end": 3.86,
              "word": "you",
              "start": 3.74,
              "speaker": "SPEAKER_01",
              "probability": 0.9921875
            },
            {
              "end": 4.1,
              "word": "about",
              "start": 3.86,
              "speaker": "SPEAKER_01",
              "probability": 0.9990234375
            },
            {
              "end": 4.48,
              "word": "AI.",
              "start": 4.1,
              "speaker": "SPEAKER_01",
              "probability": 0.966796875
            }
          ],
          "speaker": "SPEAKER_01",
          "duration": 1.5400000000000005,
          "avg_logprob": -0.17070312313735486
        },
        {
          "end": 11.66,
          "text": "It seems like this year for the entirety of the human civilization is an interesting year for the development of artificial intelligence.",
          "start": 4.72,
          "words": [
            {…

    Settings

    language
    en
03

About Whisper Diarization

TL;DRAs of September 23, 2026

Whisper Diarization is a model by Community in the Speech-to-text category. On Railwail, Whisper Diarization costs β‰ˆ $0.0063 per run.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (β‰ˆ 5 s on L40S)$0.0063 per run
GPU time (L40S)$0.00117 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3Γ— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 5.3 s

Total

$0.63

63 credits

Per run

$0.0063 Β· 0.63 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call Whisper Diarization with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
whisper-diarization
Developer
Community
Input
Audio
Output
Text
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • file_urlrequired

    URL of the audio file to transcribe and diarize

    Type: Text
    Default: –
    Allowed values: –
  • prompt

    Optional vocabulary or context hint

    Type: Text
    Default: –
    Allowed values: up to 1,000 characters
  • language

    Optional ISO-639-1 language code; auto-detected if omitted

    Type: Text
    Default: –
    Allowed values: –
  • translate
    Type: Yes/no
    Default: false
    Allowed values: –
  • num_speakers

    Optional fixed speaker count; estimated if omitted

    Type: Integer
    Default: –
    Allowed values: 1 to 50

Tags

  • replicate
  • whisper
  • stt
  • transcription
  • diarization
  • speaker-labels
  • multilingual
07

Use cases

08

Frequently asked questions

What is Whisper Diarization?

Whisper Diarization is a model by Community in the Speech-to-text category.

How much does Whisper Diarization cost on Railwail?

On Railwail, Whisper Diarization costs β‰ˆ $0.0063 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

Which settings does Whisper Diarization support?

According to its input schema, Whisper Diarization knows these parameters: file_url, prompt (up to 1,000 characters), language, translate, and num_speakers (1 to 50).

How fast is Whisper Diarization?

There are not enough measured runs of Whisper Diarization on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Whisper Diarization better than Incredibly Fast Whisper?

That depends on the task. Whisper Diarization (Community) and Incredibly Fast Whisper (Community) are both models in the Speech-to-text category. The comparison page shows their prices and specifications side by side.

Compare Whisper Diarization and Incredibly Fast Whisper
09

Comparable models

All in this category
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    β‰ˆ $0.0056/run

    11 % cheaper per unit

    Compare Whisper Diarization vs. Incredibly Fast Whisper
  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    β‰ˆ $0.0034/run

    46 % cheaper per unit

    Compare Whisper Diarization vs. Whisper
  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    β‰ˆ $0.156/run

    2,376 % more expensive per unit

    Compare Whisper Diarization vs. SeamlessM4T

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.