Whisper Diarization

Riconoscimento vocale (STT)Disponibile
di CommunityID modello: whisper-diarization

Whisper Large v3 Turbo combined with pyannote 4.0 for speaker diarization, returning who-said-what segments with timestamps. Built by Thomas Mol. Returns a clean JSON of speaker-labeled segments, handy for meeting notes, interviews, and podcasts.

Prezzo
≈ 0,0063 USD/esecuzione
Input → output
Audio → Testo
Sviluppatore
Community
Aggiornato
23 settembre 2026
01

Playground

Prova Whisper Diarization

Input e output

≈ 0,0063 USD/esecuzione
Prova Whisper Diarization

0 / 1000

URL of the audio file to transcribe and diarize

Impostazioni avanzate (2)

Optional fixed speaker count; estimated if omitted

Risultato
La trascrizione appare qui.

Questa esecuzione

circa 0,0063 USD · 0,63 crediti

0,0188 USD (1,88 crediti) sono riservati all'inizio; il tempo GPU effettivo viene fatturato.

Nuovo qui?

10 crediti gratuiti (0,10 USD) quando ti iscrivi con Google

Utilizzabile 24 ore dopo l'iscrizione, fino a 5 esecuzioni al giorno e al massimo 2 crediti per esecuzione. Altri metodi di accesso iniziano senza crediti. Sufficiente per 5 esecuzioni di questo modello.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    LLama, AI, Meta.

    Output (JSON, shortened)

    {
      "language": "en",
      "segments": [
        {
          "end": 4.48,
          "text": "Let me ask you about AI.",
          "start": 2.94,
          "words": [
            {
              "end": 3.12,
              "word": "Let",
              "start": 2.94,
              "speaker": "SPEAKER_01",
              "probability": 0.685546875
            },
            {
              "end": 3.26,
              "word": "me",
              "start": 3.12,
              "speaker": "SPEAKER_01",
              "probability": 0.9990234375
            },
            {
              "end": 3.74,
              "word": "ask",
              "start": 3.26,
              "speaker": "SPEAKER_01",
              "probability": 0.998046875
            },
            {
              "end": 3.86,
              "word": "you",
              "start": 3.74,
              "speaker": "SPEAKER_01",
              "probability": 0.9921875
            },
            {
              "end": 4.1,
              "word": "about",
              "start": 3.86,
              "speaker": "SPEAKER_01",
              "probability": 0.9990234375
            },
            {
              "end": 4.48,
              "word": "AI.",
              "start": 4.1,
              "speaker": "SPEAKER_01",
              "probability": 0.966796875
            }
          ],
          "speaker": "SPEAKER_01",
          "duration": 1.5400000000000005,
          "avg_logprob": -0.17070312313735486
        },
        {
          "end": 11.66,
          "text": "It seems like this year for the entirety of the human civilization is an interesting year for the development of artificial intelligence.",
          "start": 4.72,
          "words": [
            {…

    Settings

    language
    en
03

Informazioni su Whisper Diarization

RiassuntoA partire da 23 settembre 2026

Whisper Diarization è un modello di Community nella categoria Riconoscimento vocale (STT). Su Railwail, Whisper Diarization costa ≈ 0,0063 USD per esecuzione.

04

Prezzi

Prezzi in dollari USA. L'utilizzo viene addebitato dai crediti prepagati.
Esecuzione tipica (≈ 5 s su L40S)0,0063 USD per esecuzione
Tempo GPU (L40S)0,00117 USD per secondo GPU
  • Fatturato in base al tempo GPU effettivamente utilizzato. All'avvio, 3× il prezzo tipico viene riservato dal tuo saldo e regolato successivamente.
  • 1 credito = 0,01 USD

Calcolatore di costi

Calcolatore prezzi

s

Tipico secondo il provider: circa 5,3 s

Totale

0,63 USD

63 crediti

Per esecuzione

0,0063 USD · 0,63 crediti

Fatturato in base al tempo GPU effettivo; questo è una stima.

05

API

Chiama Whisper Diarization con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Nessun esempio API verificato

L'API pubblica passa un formato di input diverso da quello richiesto da questo modello. Usa il playground sopra.

06

Specifiche

ID modello
whisper-diarization
Sviluppatore
Community
Input
Audio
Output
Testo
Fatturazione
In base all'utilizzo (token o tempo GPU)
Voce di catalogo aggiornata
23 settembre 2026

Parametri di input

Input e impostazioni dallo schema di input del modello. L'esempio nella sezione API mostra quali di essi l'API accetta.

  • file_urlObbligatorio

    URL of the audio file to transcribe and diarize

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • prompt

    Optional vocabulary or context hint

    Tipo: Testo
    Predefinito: –
    Valori consentiti: fino a 1000 caratteri
  • language

    Optional ISO-639-1 language code; auto-detected if omitted

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • translate
    Tipo: Sì/No
    Predefinito: false
    Valori consentiti: –
  • num_speakers

    Optional fixed speaker count; estimated if omitted

    Tipo: Numero intero
    Predefinito: –
    Valori consentiti: 1 a 50

Etichette

  • replicate
  • whisper
  • stt
  • transcription
  • diarization
  • speaker-labels
  • multilingual
07

Casi d'uso

08

Domande frequenti

Cos'è Whisper Diarization?

Whisper Diarization è un modello di Community nella categoria Riconoscimento vocale (STT).

Quanto costa Whisper Diarization su Railwail?

Su Railwail, Whisper Diarization costa ≈ 0,0063 USD per esecuzione. Ti viene addebitato ciò che ogni richiesta utilizza effettivamente. L'utilizzo viene pagato con crediti prepagati; 1 credito equivale a 0,01 USD.

Quali impostazioni supporta Whisper Diarization?

Secondo il suo schema di input, Whisper Diarization conosce questi parametri: file_url, prompt (fino a 1000 caratteri), language, translate e num_speakers (1 a 50).

Quanto è veloce Whisper Diarization?

Non ci sono ancora abbastanza esecuzioni misurate di Whisper Diarization su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

Whisper Diarization è migliore di Incredibly Fast Whisper?

Dipende dall'attività. Whisper Diarization (Community) e Incredibly Fast Whisper (Community) sono entrambi modelli nella categoria Riconoscimento vocale (STT). La pagina di confronto mostra i loro prezzi e le specifiche affiancati.

Confronta Whisper Diarization e Incredibly Fast Whisper
09

Modelli comparabili

Tutti in questa categoria
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ 0,0056 USD/esecuzione

    11 % più economico per unità

    Confronta Whisper Diarization e Incredibly Fast Whisper
  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ 0,0034 USD/esecuzione

    46 % più economico per unità

    Confronta Whisper Diarization e Whisper
  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ 0,156 USD/esecuzione

    2376 % più costoso per unità

    Confronta Whisper Diarization e SeamlessM4T

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.