Whisper Large V3

Speech-to-textDeprecatedUnavailable
by OpenAIModel ID: whisper-large-v3

OpenAI's Whisper model. State-of-the-art speech recognition supporting 99+ languages.

Status
Unavailable
Input โ†’ output
Audio โ†’ Text
Developer
OpenAI
Updated
September 23, 2026

Whisper Large V3 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives

The provider is phasing this model out.

01

Comparable models

All in this category
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

02

Playground

Try Whisper Large V3

No input form

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

03

About Whisper Large V3

TL;DRAs of September 23, 2026

Whisper Large V3 is a model by OpenAI in the Speech-to-text category. Whisper Large V3 is currently not available on Railwail.

Background

About OpenAI

Founded 2015 ยท San Francisco, California, USA

OpenAI was founded in December 2015 by Sam Altman, Elon Musk, Greg Brockman, Ilya Sutskever, Wojciech Zaremba and John Schulman, restructured to capped-profit OpenAI LP in 2019. Whisper was first released as an open-weights speech recognition model in September 2022 by Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey and Ilya Sutskever, and quickly became the de-facto open ASR baseline thanks to its zero-shot multilingual quality. Whisper Large v3 was released in November 2023 alongside the GPT-4 Turbo launch as the final open checkpoint in the original Whisper line. OpenAI later released Whisper Large v3 Turbo (October 2024), a distilled faster variant. All Whisper weights remain available under MIT licence on GitHub and Hugging Face and are widely used in production by competitors.

Visit OpenAI

Architecture

Encoder-decoder Transformer for multitask speech recognition

Whisper Large v3 is a 1.55B-parameter encoder-decoder Transformer trained for multitask audio understanding. Input is a 30-second log-mel spectrogram with 128 mel bins (up from 80 in v2) processed by a 32-layer audio encoder; the decoder is a 32-layer Transformer that emits special task tokens for language identification, transcription, translation-to-English and timestamp prediction. Whisper was trained on 5 million hours of audio in total: 680k hours of weakly supervised multilingual web audio for v1/v2 plus an additional 4 million hours of pseudo-labelled audio generated by Whisper Large v2 for v3, yielding approximately 1 million hours of weakly supervised and 4 million hours of pseudo-labelled data. The model covers 99 languages, with very strong performance on English long-form (~5% WER on LibriSpeech test-clean) and significantly improved low-resource coverage. Audio longer than 30 seconds is processed with a sliding window. The model is released under MIT licence.

Parameters
1.55B (Large v3)
Context
30 tokens

Capabilities

  • 99-language transcription with automatic language detection
  • Translation of any supported language directly to English text
  • Word-level and segment-level timestamps
  • 30-second context window with sliding window for long-form audio
  • Open weights under MIT licence, runs on a single 16 GB GPU in fp16
  • Strong robustness to accents, background noise and technical jargon
  • Best for: open-source ASR pipelines, research baselines, on-premise transcription

Training & license

5 million hours total: ~680k hours of weakly supervised multilingual audio scraped from the public web (subtitles aligned with audio) plus ~4 million hours of pseudo-labelled audio generated by Whisper Large v2. 17% of the supervised set is non-English speech.

License: MIT licence for code and weights; commercial use permitted.

Safety testing: OpenAI published bias and hallucination analyses; the paper documents over-confident hallucination on silent or non-speech audio. No formal red-team report beyond the original paper.

Known limitations

  • 30-second hard window requires chunking for long audio
  • Known to hallucinate transcripts on silent / music-only segments
  • WER on low-resource languages still much higher than English
  • No native diarisation (must be combined with pyannote)
  • Slow on CPU; requires GPU for real-time
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call Whisper Large V3 with your Railwail API key. Use this model ID in the request:

Currently unavailable

The model has no verified price or is deactivated; API calls are refused.

06

Specifications

Model ID
whisper-large-v3
Developer
OpenAI
Input
Audio
Output
Text
Output formats
JSON, SRT, VTT
Lifecycle
Deprecated
Model size
1.55B (Large v3)
License
MIT licence for code and weights; commercial use permitted.
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • filerequired
    Type: File
    Default: โ€“
    Allowed values: audio/*
  • language
    Type: Text
    Default: โ€“
    Allowed values: โ€“

Tags

  • multilingual
  • popular
07

Example prompts

Examples from the Railwail catalog. They were not generated live on this page.
  • Meeting Notes

    Show example answer

    Good morning everyone. Let's start with the sprint review. The authentication module shipped on Friday and we've had zero critical bugs reported so far. Sarah, can you walk us through the performance metrics? We're seeing a forty percent reduction in login latency which is well above our target.

  • Interview Transcript

    Show example answer

    So tell me about your experience with distributed systems. I spent three years at a fintech startup building event-driven microservices using Kafka and Kubernetes. The biggest challenge was handling exactly-once delivery semantics across our payment processing pipeline. We ended up implementing an idempotency layer that reduced duplicate transactions by ninety-nine point nine percent.

08

Use cases

What it is used for

  • Open-source transcription pipelines
  • Academic ASR research baselines
  • Multilingual podcast transcription
  • Translation of foreign-language audio to English
  • On-premise speech recognition for privacy-sensitive workloads
09

Frequently asked questions

What is Whisper Large V3?

Whisper Large V3 is a model by OpenAI in the Speech-to-text category. It is listed on Railwail but cannot be run at the moment.

How much does Whisper Large V3 cost on Railwail?

Whisper Large V3 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

Which settings does Whisper Large V3 support?

According to its input schema, Whisper Large V3 knows these parameters: file (audio/*) and language.

How fast is Whisper Large V3?

There are not enough measured runs of Whisper Large V3 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Whisper Large V3 better than Incredibly Fast Whisper?

That depends on the task. Whisper Large V3 (OpenAI) and Incredibly Fast Whisper (Community) are both models in the Speech-to-text category. The comparison page shows their prices and specifications side by side.

Compare Whisper Large V3 and Incredibly Fast Whisper

Can I use Whisper Large V3 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.