Whisper Large v3 Turbo

Rozpoznawanie mowyNiedostępne
od OpenAIID modelu: whisper-large-v3-turbo

OpenAI's distilled Whisper Large v3. ~216x realtime, 99+ languages, MIT-licensed weights.

Status
Niedostępne
Wejście → Wyjście
Audio → Tekst
Deweloper
OpenAI
Zaktualizowano
23 września 2026

Whisper Large v3 Turbo jest obecnie niedostępny

Obecnie niedostępne: ten model został dezaktywowany.

Możesz nadal przeczytać szczegóły na tej stronie. Wybierz jedną z dostępnych alternatyw poniżej, aby od razu uruchomić porównywalny model.

Przejdź do alternatyw
01

Porównywalne modele

Wszystkie w tej kategorii
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ 0,0056 USD/uruchomienie

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ 0,0034 USD/uruchomienie

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ 0,156 USD/uruchomienie

02

Playground

Spróbuj Whisper Large v3 Turbo

Brak formularza wejściowego

Niedostępny

Obecnie niedostępne: ten model został dezaktywowany.

Plac zabaw jest wyłączony. Porównywalne modele znajdziesz w tej samej kategorii: Przeglądaj alternatywy

03

O Whisper Large v3 Turbo

Krótko mówiącStan na 23 września 2026

Whisper Large v3 Turbo to model opracowany przez OpenAI w kategorii Rozpoznawanie mowy. Whisper Large v3 Turbo nie jest obecnie dostępny w serwisie Railwail.

Tło

O OpenAI

Założona 2015 · San Francisco, California, USA

OpenAI was founded in December 2015 by Sam Altman, Elon Musk, Greg Brockman, Ilya Sutskever, Wojciech Zaremba and John Schulman, restructured to capped-profit OpenAI LP in 2019. Whisper Large v3 Turbo was released in October 2024 as a distilled fast variant of Whisper Large v3, designed to deliver approximately 8x faster inference at near-identical accuracy by reducing the decoder depth from 32 to 4 layers. The release was led by the original Whisper authors (Alec Radford, Jong Wook Kim, Tao Xu) and remained under the MIT licence. Turbo was distributed via GitHub, the Hugging Face Hub and the OpenAI Whisper API as the new default model where supported, replacing many in-production deployments of Whisper Large v3 within weeks of launch.

Odwiedź OpenAI

Architektura

Distilled encoder-decoder Transformer (4-layer decoder) for speech recognition

Whisper Large v3 Turbo is a distilled variant of Whisper Large v3 that keeps the same 32-layer audio encoder and 128-mel front-end but shrinks the decoder from 32 to just 4 Transformer layers, taking the total parameter count from 1.55B to 809M. The smaller decoder gives roughly 8x faster inference on long-form audio (and 4-5x faster on short clips) at a WER cost of approximately 0.5-1 percentage points on most benchmarks. The model was distilled on the same multilingual corpus as Large v3 (5 million hours total, of which 4 million are pseudo-labelled) with knowledge-distillation losses from the Large v3 teacher. Translation-to-English capability was deliberately removed to focus capacity on transcription quality. The 30-second sliding window, 99-language coverage and special task tokens are unchanged. Turbo runs in real-time on consumer GPUs (RTX 3060) and at 3-4x real-time on Apple Silicon CPUs via whisper.cpp.

Parametry
809M
Kontekst
30 tokenów

Możliwości

  • 8x faster long-form transcription than Whisper Large v3
  • 99-language transcription with automatic language detection
  • Word-level timestamps preserved
  • Runs in real-time on a single consumer GPU (RTX 3060 / M2 Pro)
  • Half the memory footprint of Large v3 (809M vs 1.55B)
  • Open weights under MIT licence
  • Drop-in replacement for Large v3 in most pipelines
  • Best for: production ASR on commodity hardware, on-premise transcription, batch processing

Trening i licencja

Distilled from Whisper Large v3 on the same 5-million-hour multilingual audio corpus with knowledge-distillation losses. Translation-to-English data was excluded.

Licencja: MIT licence for code and weights; commercial use permitted.

Testy bezpieczeństwa: Inherits all hallucination and bias caveats from Whisper Large v3; no separate red-team report.

Znane ograniczenia

  • No translation-to-English mode (transcription only)
  • WER 0.5-1 pp worse than Large v3 on average
  • Same 30-second hard window requires chunking
  • Same hallucination behaviour on silent / music-only audio
  • No native diarisation
04

Ceny

Obecnie niedostępne: ten model został dezaktywowany. Dla tego modelu nie ma ceny w tej chwili, dlatego nie można go uruchomić.

05

API

Wywołaj Whisper Large v3 Turbo za pomocą klucza API Railwail. Użyj tego ID modelu w żądaniu:
whisper-large-v3-turboDokumentacja APIUzyskaj klucz API

Obecnie niedostępne

Model nie ma zweryfikowanej ceny lub jest wyłączony; wywołania API są odrzucane.

06

Specyfikacje

ID modelu
whisper-large-v3-turbo
Deweloper
OpenAI
Wejście
Audio
Wyjście
Tekst
Formaty wyjścia
JSON, SRT, VTT
Cykl życia
Niedostępne
Rozmiar modelu
809M
Licencja
MIT licence for code and weights; commercial use permitted.
Wpis w katalogu zaktualizowany
23 września 2026

Parametry wejściowe

Dane wejściowe i ustawienia ze schematu wejściowego modelu. Przykład w sekcji API pokazuje, które z nich API akceptuje.

  • filewymagane

    URL or upload path to audio file

    Typ: Tekst
    Domyślnie: –
    Dozwolone wartości: –
  • prompt

    Optional context to guide transcription

    Typ: Tekst
    Domyślnie: –
    Dozwolone wartości: do 1000 znaków
  • language

    Optional ISO-639-1 language code (e.g. en, de, fr)

    Typ: Tekst
    Domyślnie: –
    Dozwolone wartości: –
  • temperature
    Typ: Liczba
    Domyślnie: 0
    Dozwolone wartości: 0 do 1
  • response_format
    Typ: Wybór
    Domyślnie: json
    Dozwolone wartości: json, text, srt lub vtt

Tagi

  • openai
  • whisper
  • stt
  • transcription
  • open-weights
  • multilingual
  • per-minute
07

Przypadki użycia

Do czego się go używa

  • Real-time transcription on consumer GPUs
  • Batch transcription of large podcast / lecture archives
  • On-device transcription via whisper.cpp
  • Cost-sensitive production ASR pipelines
  • Edge deployments where bandwidth is limited
08

Często zadawane pytania

Co to jest Whisper Large v3 Turbo?

Whisper Large v3 Turbo to model opracowany przez OpenAI w kategorii Rozpoznawanie mowy. Jest wymieniony w katalogu Railwail, ale nie może być uruchomiony w tej chwili.

Ile kosztuje Whisper Large v3 Turbo w serwisie Railwail?

Whisper Large v3 Turbo nie może być uruchomiony w serwisie Railwail w tej chwili, dlatego nie ma aktualnej ceny. Dostępne alternatywy z cenami są wymienione poniżej na tej stronie.

Jakie ustawienia obsługuje Whisper Large v3 Turbo?

Zgodnie ze schematem wejściowym, Whisper Large v3 Turbo obsługuje te parametry: file, prompt (do 1000 znaków), language, temperature (0 do 1) i response_format (json, text, srt lub vtt).

Jak szybki jest Whisper Large v3 Turbo?

Dla Whisper Large v3 Turbo jest jeszcze zbyt mało zmierzonych przebiegów w serwisie Railwail, aby podać czas przebiegu. Zależy to od wejścia, ustawień i obciążenia u dostawcy.

Czy Whisper Large v3 Turbo jest lepszy niż Incredibly Fast Whisper?

To zależy od zadania. Whisper Large v3 Turbo (OpenAI) i Incredibly Fast Whisper (Community) to oba modele z kategorii Rozpoznawanie mowy. Strona porównania pokazuje ich ceny i specyfikacje obok siebie.

Porównaj Whisper Large v3 Turbo i Incredibly Fast Whisper

Czy mogę używać Whisper Large v3 Turbo teraz?

Obecnie niedostępne: ten model został dezaktywowany. Strona pozostaje online; dostępne alternatywy z tej samej kategorii są wymienione poniżej.

Wszystkie modele przez jedno API

Jeden klucz API dla każdego modelu na Railwail. Opłaty pobierane są z przedpłaconych kredytów, 1 kredyt = 0,01 USD.