Deepgram Nova-3

Speech-to-textUnavailable
by OtherModel ID: nova-3

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

Status
Unavailable
Input โ†’ output
Audio โ†’ Text
Developer
Other
Updated
23 September 2026

Deepgram Nova-3 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

02

Playground

Try Deepgram Nova-3

Input & output

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Deepgram Nova-3

0 / 1,000

Optional context to guide transcription

URL or upload path to audio file

Optional ISO-639-1 language code (e.g. en, de, fr)

Advanced settings (2)
Output
The transcript appears here.

This run

No price โ€“ currently unavailable.

New here?

10 free credits (US$0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Deepgram Nova-3

TL;DRAs of 23 September 2026

Deepgram Nova-3 is a model by Other in the Speech-to-text category. Deepgram Nova-3 is currently not available on Railwail.

Background

About Deepgram

Founded 2015 ยท San Francisco, California, USA

Deepgram was founded in 2015 by Scott Stephenson (CEO) and Noah Shutty after they finished PhDs in dark-matter physics at the University of Michigan and built a Hadoop pipeline to search audio archives. The company pivoted to commercial speech recognition in 2018 and became one of the first speech vendors to ship a fully end-to-end deep-learning ASR pipeline rather than the classical Kaldi stack. Deepgram has raised over $86M in venture funding from Tiger Global, Y Combinator, Wing, Madrona and NVIDIA's NVentures, with a 2024 Series C at a reported $1B+ valuation. The Nova model family launched in 2023 (Nova-1), followed by Nova-2 (2024) and Nova-3 (January 2025), positioning Deepgram as the fastest commercial ASR provider with sub-300 ms streaming and self-hostable deployments.

Visit Deepgram

Architecture

End-to-end encoder-decoder speech-to-text with self-supervised pretraining

Deepgram Nova-3 is an end-to-end deep-learning automatic speech recognition model trained on what Deepgram describes as the largest commercial training corpus among production ASR providers, including 50,000+ hours of curated multi-domain audio plus self-supervised pretraining on hundreds of thousands of hours of unlabelled speech. Nova-3 introduced a unified multilingual model that natively handles English, Spanish, French, German, Italian, Portuguese, Hindi, Japanese, Mandarin, Korean and 10+ more, with on-the-fly code-switching inside a single audio file. New capabilities include keyword prompting (custom vocabulary at request time without retraining), Self-Hosted edition for HIPAA/SOC2 deployments, and improved noise robustness on call-centre and phone audio. The streaming variant runs at sub-300 ms latency over WebSocket, while the pre-recorded variant supports up to 4 hours per file. Outputs include diarisation, punctuation, smart formatting, language detection and topic/intent metadata.

Parameters
Undisclosed (multi-billion)
Context
14,400 tokens

Capabilities

  • Real-time streaming ASR at sub-300 ms latency
  • Multilingual single model with code-switching across 36+ languages
  • Keyword prompting for runtime custom vocabulary
  • Speaker diarisation, punctuation and smart formatting
  • Phone-call optimised (8 kHz, low-bandwidth, noisy)
  • Self-Hosted deployment for HIPAA / SOC2 / on-premises
  • Up to 4 hours of audio per pre-recorded request
  • Best for: call centres, real-time meeting transcription, voice agents, compliance recording

Training & license

Self-supervised pretraining on hundreds of thousands of hours of unlabelled audio plus supervised training on 50,000+ hours of curated, labelled multilingual speech. Data is sourced from licensed corpora, customer opt-in audio and publicly available datasets.

License: Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.

Safety testing: SOC 2 Type II and HIPAA compliant; no model-output watermarking required for transcripts; bias testing reported in the Nova-3 launch blog.

Known limitations

  • Word Error Rate worse than Whisper Large v3 on some long-form audiobook tasks
  • Streaming diarisation can mislabel speakers in fast cross-talk
  • Self-Hosted edition has high hardware requirements
  • Code-switching quality varies between language pairs
  • Closed weights for hosted API
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call Deepgram Nova-3 with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
nova-3
Developer
Other
Input
Audio
Output
Text
Output formats
JSON, SRT, VTT
Model size
Undisclosed (multi-billion)
License
Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.
Catalog entry updated
23 September 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • filerequired

    URL or upload path to audio file

    Type: Text
    Default: โ€“
    Allowed values: โ€“
  • prompt

    Optional context to guide transcription

    Type: Text
    Default: โ€“
    Allowed values: up to 1,000 characters
  • language

    Optional ISO-639-1 language code (e.g. en, de, fr)

    Type: Text
    Default: โ€“
    Allowed values: โ€“
  • temperature
    Type: Number
    Default: 0
    Allowed values: 0 to 1
  • response_format
    Type: Choice
    Default: json
    Allowed values: json, text, srt or vtt

Tags

  • deepgram
  • stt
  • transcription
  • realtime
  • multilingual
  • per-minute
07

Use cases

What it is used for

  • Real-time call-centre transcription and analytics
  • Voice agent transcript layer
  • Live closed captioning
  • Compliance recording for finance and healthcare
  • Meeting transcription with diarisation
08

Frequently asked questions

What is Deepgram Nova-3?

Deepgram Nova-3 is a model by Other in the Speech-to-text category. It is listed on Railwail but cannot be run at the moment.

How much does Deepgram Nova-3 cost on Railwail?

Deepgram Nova-3 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

Which settings does Deepgram Nova-3 support?

According to its input schema, Deepgram Nova-3 knows these parameters: file, prompt (up to 1,000 characters), language, temperature (0 to 1) and response_format (json, text, srt or vtt).

How fast is Deepgram Nova-3?

There are not enough measured runs of Deepgram Nova-3 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Deepgram Nova-3 better than Incredibly Fast Whisper?

That depends on the task. Deepgram Nova-3 (Other) and Incredibly Fast Whisper (Community) are both models in the Speech-to-text category. The comparison page shows their prices and specifications side by side.

Compare Deepgram Nova-3 and Incredibly Fast Whisper

Can I use Deepgram Nova-3 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.