WhisperX

Riconoscimento vocale (STT)Disponibile
di CommunityID modello: whisperx

WhisperX (Large v3) with forced alignment for accurate word-level timestamps plus optional speaker diarization. Uses VAD to cut long files into segments and a wav2vec2 aligner to pin each word to its exact time. Useful for subtitles and per-speaker transcripts.

Prezzo
≈ 0,0241 USD/esecuzione
Input → output
Audio → Testo
Sviluppatore
Community
Aggiornato
23 settembre 2026
01

Playground

Prova WhisperX

Input e output

≈ 0,0241 USD/esecuzione
Prova WhisperX

URL or upload of the audio file to transcribe

Impostazioni avanzate (3)
Risultato
La trascrizione appare qui.

Questa esecuzione

circa 0,0241 USD · 2,41 crediti

0,0721 USD (7,21 crediti) sono riservati all'inizio; il tempo GPU effettivo viene fatturato.

Per account senza acquisti precedenti: le esecuzioni oltre 2 crediti richiedono una ricarica.

Nuovo qui?

10 crediti gratuiti (0,10 USD) quando ti iscrivi con Google

Utilizzabile 24 ore dopo l'iscrizione, fino a 5 esecuzioni al giorno e al massimo 2 crediti per esecuzione. Altri metodi di accesso iniziano senza crediti.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Output (JSON, shortened)

    {
      "segments": [
        {
          "end": 30.811,
          "text": "The little tales they tell are false. The door was barred, locked and bolted as well. Ripe pears are fit for a queen's table. A big wet stain was on the round carpet. The kite dipped and swayed but stayed aloft. The pleasant hours fly by much too soon. The room was crowded with a mild wob.",
          "start": 2.585
        },
        {
          "end": 48.592,
          "text": "The room was crowded with a wild mob. This strong arm shall shield your honor. She blushed when he gave her a white orchid. The beetle droned in the hot June sun.",
          "start": 33.029
        }
      ],
      "detected_language": "en"
    }
03

Informazioni su WhisperX

RiassuntoA partire da 23 settembre 2026

WhisperX è un modello di Community nella categoria Riconoscimento vocale (STT). Su Railwail, WhisperX costa ≈ 0,0241 USD per esecuzione.

04

Prezzi

Prezzi in dollari USA. L'utilizzo viene addebitato dai crediti prepagati.
Esecuzione tipica (≈ 14 s su A100 (80GB))0,0241 USD per esecuzione
Tempo GPU (A100 (80GB))0,00168 USD per secondo GPU
  • Fatturato in base al tempo GPU effettivamente utilizzato. All'avvio, 3× il prezzo tipico viene riservato dal tuo saldo e regolato successivamente.
  • 1 credito = 0,01 USD

Calcolatore di costi

Calcolatore prezzi

s

Tipico secondo il provider: circa 14,3 s

Totale

2,41 USD

241 crediti

Per esecuzione

0,0241 USD · 2,41 crediti

Fatturato in base al tempo GPU effettivo; questo è una stima.

05

API

Chiama WhisperX con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Nessun esempio API verificato

L'API pubblica passa un formato di input diverso da quello richiesto da questo modello. Usa il playground sopra.

06

Specifiche

ID modello
whisperx
Sviluppatore
Community
Input
Audio
Output
Testo
Fatturazione
In base all'utilizzo (token o tempo GPU)
Voce di catalogo aggiornata
23 settembre 2026

Parametri di input

Input e impostazioni dallo schema di input del modello. L'esempio nella sezione API mostra quali di essi l'API accetta.

  • audio_fileObbligatorio

    URL or upload of the audio file to transcribe

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • language

    Optional ISO-639-1 language code; auto-detected if omitted

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • batch_size
    Tipo: Numero intero
    Predefinito: 64
    Valori consentiti: 1 a 64
  • diarization
    Tipo: Sì/No
    Predefinito: false
    Valori consentiti: –
  • align_output
    Tipo: Sì/No
    Predefinito: true
    Valori consentiti: –

Etichette

  • replicate
  • whisperx
  • stt
  • transcription
  • diarization
  • word-timestamps
  • multilingual
07

Casi d'uso

08

Domande frequenti

Cos'è WhisperX?

WhisperX è un modello di Community nella categoria Riconoscimento vocale (STT).

Quanto costa WhisperX su Railwail?

Su Railwail, WhisperX costa ≈ 0,0241 USD per esecuzione. Ti viene addebitato ciò che ogni richiesta utilizza effettivamente. L'utilizzo viene pagato con crediti prepagati; 1 credito equivale a 0,01 USD.

Quali impostazioni supporta WhisperX?

Secondo il suo schema di input, WhisperX conosce questi parametri: audio_file, language, batch_size (1 a 64), diarization e align_output.

Quanto è veloce WhisperX?

Non ci sono ancora abbastanza esecuzioni misurate di WhisperX su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

WhisperX è migliore di Incredibly Fast Whisper?

Dipende dall'attività. WhisperX (Community) e Incredibly Fast Whisper (Community) sono entrambi modelli nella categoria Riconoscimento vocale (STT). La pagina di confronto mostra i loro prezzi e le specifiche affiancati.

Confronta WhisperX e Incredibly Fast Whisper
09

Modelli comparabili

Tutti in questa categoria
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ 0,0056 USD/esecuzione

    77 % più economico per unità

    Confronta WhisperX e Incredibly Fast Whisper
  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ 0,0034 USD/esecuzione

    86 % più economico per unità

    Confronta WhisperX e Whisper
  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ 0,156 USD/esecuzione

    547 % più costoso per unità

    Confronta WhisperX e SeamlessM4T

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.