Whisper Large v3 Turbo

Conversão de fala em textoIndisponível
por OpenAIID do modelo: whisper-large-v3-turbo

OpenAI's distilled Whisper Large v3. ~216x realtime, 99+ languages, MIT-licensed weights.

Status
Indisponível
Entrada → Saída
Áudio → Texto
Desenvolvedor
OpenAI
Atualizado
23 de setembro de 2026

Whisper Large v3 Turbo não está disponível no momento

Atualmente indisponível: este modelo foi desativado.

Você ainda pode ler os detalhes nesta página. Escolha uma das alternativas disponíveis abaixo para executar um modelo comparável imediatamente.

Ir para alternativas
01

Modelos comparáveis

Todos nesta categoria
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ US$ 0,0056/execução

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ US$ 0,0034/execução

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ US$ 0,156/execução

02

Playground

Experimentar Whisper Large v3 Turbo

Sem formulário de entrada

Indisponível no momento

Atualmente indisponível: este modelo foi desativado.

O playground está desativado. Encontre modelos comparáveis na mesma categoria: Ver alternativas

03

Sobre Whisper Large v3 Turbo

ResumoA partir de 23 de setembro de 2026

Whisper Large v3 Turbo é um modelo de OpenAI na categoria Conversão de fala em texto. Whisper Large v3 Turbo não está disponível no Railwail no momento.

Fundo

Sobre OpenAI

Fundado em 2015 · San Francisco, California, USA

OpenAI was founded in December 2015 by Sam Altman, Elon Musk, Greg Brockman, Ilya Sutskever, Wojciech Zaremba and John Schulman, restructured to capped-profit OpenAI LP in 2019. Whisper Large v3 Turbo was released in October 2024 as a distilled fast variant of Whisper Large v3, designed to deliver approximately 8x faster inference at near-identical accuracy by reducing the decoder depth from 32 to 4 layers. The release was led by the original Whisper authors (Alec Radford, Jong Wook Kim, Tao Xu) and remained under the MIT licence. Turbo was distributed via GitHub, the Hugging Face Hub and the OpenAI Whisper API as the new default model where supported, replacing many in-production deployments of Whisper Large v3 within weeks of launch.

Visite OpenAI

Arquitetura

Distilled encoder-decoder Transformer (4-layer decoder) for speech recognition

Whisper Large v3 Turbo is a distilled variant of Whisper Large v3 that keeps the same 32-layer audio encoder and 128-mel front-end but shrinks the decoder from 32 to just 4 Transformer layers, taking the total parameter count from 1.55B to 809M. The smaller decoder gives roughly 8x faster inference on long-form audio (and 4-5x faster on short clips) at a WER cost of approximately 0.5-1 percentage points on most benchmarks. The model was distilled on the same multilingual corpus as Large v3 (5 million hours total, of which 4 million are pseudo-labelled) with knowledge-distillation losses from the Large v3 teacher. Translation-to-English capability was deliberately removed to focus capacity on transcription quality. The 30-second sliding window, 99-language coverage and special task tokens are unchanged. Turbo runs in real-time on consumer GPUs (RTX 3060) and at 3-4x real-time on Apple Silicon CPUs via whisper.cpp.

Parâmetros
809M
Contexto
30 tokens

Capacidades

  • 8x faster long-form transcription than Whisper Large v3
  • 99-language transcription with automatic language detection
  • Word-level timestamps preserved
  • Runs in real-time on a single consumer GPU (RTX 3060 / M2 Pro)
  • Half the memory footprint of Large v3 (809M vs 1.55B)
  • Open weights under MIT licence
  • Drop-in replacement for Large v3 in most pipelines
  • Best for: production ASR on commodity hardware, on-premise transcription, batch processing

Treinamento & licença

Distilled from Whisper Large v3 on the same 5-million-hour multilingual audio corpus with knowledge-distillation losses. Translation-to-English data was excluded.

Licença: MIT licence for code and weights; commercial use permitted.

Testes de segurança: Inherits all hallucination and bias caveats from Whisper Large v3; no separate red-team report.

Limitações conhecidas

  • No translation-to-English mode (transcription only)
  • WER 0.5-1 pp worse than Large v3 on average
  • Same 30-second hard window requires chunking
  • Same hallucination behaviour on silent / music-only audio
  • No native diarisation
04

Preços

Atualmente indisponível: este modelo foi desativado. Não há preço para este modelo no momento, portanto não pode ser executado.

05

API

Chame Whisper Large v3 Turbo com sua chave de API Railwail. Use este ID de modelo na solicitação:

Indisponível no momento

O modelo não tem preço verificado ou está desativado; chamadas de API são recusadas.

06

Especificações

ID do modelo
whisper-large-v3-turbo
Desenvolvedor
OpenAI
Entrada
Áudio
Saída
Texto
Formatos de saída
JSON, SRT, VTT
Ciclo de vida
Indisponível
Tamanho do modelo
809M
Licença
MIT licence for code and weights; commercial use permitted.
Entrada do catálogo atualizada
23 de setembro de 2026

Parâmetros de entrada

Entradas e configurações do esquema de entrada do modelo. O exemplo na seção API mostra quais delas a API aceita.

  • fileobrigatório

    URL or upload path to audio file

    Tipo: Texto
    Padrão: –
    Valores permitidos: –
  • prompt

    Optional context to guide transcription

    Tipo: Texto
    Padrão: –
    Valores permitidos: até 1.000 caracteres
  • language

    Optional ISO-639-1 language code (e.g. en, de, fr)

    Tipo: Texto
    Padrão: –
    Valores permitidos: –
  • temperature
    Tipo: Número
    Padrão: 0
    Valores permitidos: 0 a 1
  • response_format
    Tipo: Escolha
    Padrão: json
    Valores permitidos: json, text, srt ou vtt

Etiquetas

  • openai
  • whisper
  • stt
  • transcription
  • open-weights
  • multilingual
  • per-minute
07

Casos de uso

Para que é utilizado

  • Real-time transcription on consumer GPUs
  • Batch transcription of large podcast / lecture archives
  • On-device transcription via whisper.cpp
  • Cost-sensitive production ASR pipelines
  • Edge deployments where bandwidth is limited
08

Perguntas frequentes

O que é Whisper Large v3 Turbo?

Whisper Large v3 Turbo é um modelo de OpenAI na categoria Conversão de fala em texto. Está listado no Railwail, mas não pode ser executado no momento.

Quanto custa Whisper Large v3 Turbo no Railwail?

Whisper Large v3 Turbo não pode ser executado no Railwail no momento, portanto não há preço atual. Alternativas disponíveis com preços estão listadas mais abaixo nesta página.

Quais configurações Whisper Large v3 Turbo suporta?

De acordo com seu esquema de entrada, Whisper Large v3 Turbo conhece estes parâmetros: file, prompt (até 1.000 caracteres), language, temperature (0 a 1) e response_format (json, text, srt ou vtt).

Qual é a velocidade de Whisper Large v3 Turbo?

Ainda não há execuções medidas suficientes de Whisper Large v3 Turbo no Railwail para indicar um tempo de execução. Depende da entrada, das configurações e da carga no provedor.

Whisper Large v3 Turbo é melhor que Incredibly Fast Whisper?

Depende da tarefa. Whisper Large v3 Turbo (OpenAI) e Incredibly Fast Whisper (Community) são ambos modelos na categoria Conversão de fala em texto. A página de comparação mostra seus preços e especificações lado a lado.

Comparar Whisper Large v3 Turbo e Incredibly Fast Whisper

Posso usar Whisper Large v3 Turbo agora?

Atualmente indisponível: este modelo foi desativado. A página permanece online; alternativas disponíveis da mesma categoria estão listadas mais abaixo.

Todos os modelos através de uma API

Uma chave API para todos os modelos no Railwail. O uso é cobrado a partir de créditos pré-pagos, 1 crédito = US$ 0,01.