Deepgram Nova-3

Talestemme-til-tekst (STT)Ikke tilgængelig
af OtherModell-ID: nova-3

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

Status
Ikke tilgængelig
Input → output
Lyd → Tekst
Udvikler
Other
Opdateret
23. september 2026

Deepgram Nova-3 er i øjeblikket utilgængelig

Du kan stadig læse detaljerne på denne side. Vælg en af de tilgængelige alternativer nedenfor for at køre en sammenlignelig model med det samme.

Gå til alternativer
01

Sammenlignelige modeller

Alle i denne kategori
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

02

Playground

Prøv Deepgram Nova-3

Input & resultat

Derzeit nicht verfügbar

Ikke tilgængelig i øjeblikket.

Playground'en er deaktiveret. Du finder sammenlignelige modeller i samme kategori: Se alternativer

Prøv Deepgram Nova-3

0 / 1.000

URL or upload path to audio file

Avancerede indstillinger (2)
Resultat
Transskriptionen vises her.

Denne kørsel

Ingen pris – ikke tilgængelig i øjeblikket.

Ny her?

10 gratis credits (0,10 US$) når du tilmelder dig med Google

Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits.

03

Om Deepgram Nova-3

Kort sagtFra 23. september 2026

Deepgram Nova-3 er en model af Other i kategorien Talestemme-til-tekst (STT). Deepgram Nova-3 er i øjeblikket ikke tilgængelig på Railwail.

Baggrund

Om Deepgram

Grundlagt 2015 · San Francisco, California, USA

Deepgram was founded in 2015 by Scott Stephenson (CEO) and Noah Shutty after they finished PhDs in dark-matter physics at the University of Michigan and built a Hadoop pipeline to search audio archives. The company pivoted to commercial speech recognition in 2018 and became one of the first speech vendors to ship a fully end-to-end deep-learning ASR pipeline rather than the classical Kaldi stack. Deepgram has raised over $86M in venture funding from Tiger Global, Y Combinator, Wing, Madrona and NVIDIA's NVentures, with a 2024 Series C at a reported $1B+ valuation. The Nova model family launched in 2023 (Nova-1), followed by Nova-2 (2024) and Nova-3 (January 2025), positioning Deepgram as the fastest commercial ASR provider with sub-300 ms streaming and self-hostable deployments.

Besøg Deepgram

Arkitektur

End-to-end encoder-decoder speech-to-text with self-supervised pretraining

Deepgram Nova-3 is an end-to-end deep-learning automatic speech recognition model trained on what Deepgram describes as the largest commercial training corpus among production ASR providers, including 50,000+ hours of curated multi-domain audio plus self-supervised pretraining on hundreds of thousands of hours of unlabelled speech. Nova-3 introduced a unified multilingual model that natively handles English, Spanish, French, German, Italian, Portuguese, Hindi, Japanese, Mandarin, Korean and 10+ more, with on-the-fly code-switching inside a single audio file. New capabilities include keyword prompting (custom vocabulary at request time without retraining), Self-Hosted edition for HIPAA/SOC2 deployments, and improved noise robustness on call-centre and phone audio. The streaming variant runs at sub-300 ms latency over WebSocket, while the pre-recorded variant supports up to 4 hours per file. Outputs include diarisation, punctuation, smart formatting, language detection and topic/intent metadata.

Parametre
Undisclosed (multi-billion)
Kontekst
14.400 tokens

Funktioner

  • Real-time streaming ASR at sub-300 ms latency
  • Multilingual single model with code-switching across 36+ languages
  • Keyword prompting for runtime custom vocabulary
  • Speaker diarisation, punctuation and smart formatting
  • Phone-call optimised (8 kHz, low-bandwidth, noisy)
  • Self-Hosted deployment for HIPAA / SOC2 / on-premises
  • Up to 4 hours of audio per pre-recorded request
  • Best for: call centres, real-time meeting transcription, voice agents, compliance recording

Træning og licens

Self-supervised pretraining on hundreds of thousands of hours of unlabelled audio plus supervised training on 50,000+ hours of curated, labelled multilingual speech. Data is sourced from licensed corpora, customer opt-in audio and publicly available datasets.

Licens: Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.

Sikkerhedstests: SOC 2 Type II and HIPAA compliant; no model-output watermarking required for transcripts; bias testing reported in the Nova-3 launch blog.

Kendte begrænsninger

  • Word Error Rate worse than Whisper Large v3 on some long-form audiobook tasks
  • Streaming diarisation can mislabel speakers in fast cross-talk
  • Self-Hosted edition has high hardware requirements
  • Code-switching quality varies between language pairs
  • Closed weights for hosted API
04

Priser

Ikke tilgængelig i øjeblikket. Der er i øjeblikket ingen pris for denne model, så den kan ikke køres.

05

API

Kald Deepgram Nova-3 med din Railwail API-nøgle. Brug dette model-ID i anmodningen:

Intet bekræftet API-eksempel

Den offentlige API sender et andet inputformat end det, som denne model har brug for. Brug playground ovenfor.

06

Specifikationer

Model-ID
nova-3
Udvikler
Other
Input
Lyd
Output
Tekst
Outputformater
JSON, SRT, VTT
Modelstørrelse
Undisclosed (multi-billion)
Licens
Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.
Katalogelement opdateret
23. september 2026

Inputparametre

Inputs og indstillinger fra modellens inputskema. Eksemplet i API-afsnittet viser, hvilke af dem API'en accepterer.

  • filepåkrævet

    URL or upload path to audio file

    Type: Tekst
    Standard: –
    Tilladte værdier: –
  • prompt

    Optional context to guide transcription

    Type: Tekst
    Standard: –
    Tilladte værdier: op til 1.000 tegn
  • language

    Optional ISO-639-1 language code (e.g. en, de, fr)

    Type: Tekst
    Standard: –
    Tilladte værdier: –
  • temperature
    Type: Tal
    Standard: 0
    Tilladte værdier: 0 til 1
  • response_format
    Type: Valg
    Standard: json
    Tilladte værdier: json, text, srt eller vtt

Tags

  • deepgram
  • stt
  • transcription
  • realtime
  • multilingual
  • per-minute
07

Anvendelsestilfælde

Hvad det bruges til

  • Real-time call-centre transcription and analytics
  • Voice agent transcript layer
  • Live closed captioning
  • Compliance recording for finance and healthcare
  • Meeting transcription with diarisation
08

Ofte stillede spørgsmål

Hvad er Deepgram Nova-3?

Deepgram Nova-3 er en model fra Other i kategorien Talestemme-til-tekst (STT). Den er opført på Railwail, men kan ikke køres i øjeblikket.

Hvad koster Deepgram Nova-3 på Railwail?

Deepgram Nova-3 kan ikke køres på Railwail i øjeblikket, så der er ingen aktuel pris. Tilgængelige alternativer med priser er angivet længere nede på denne side.

Hvilke indstillinger understøtter Deepgram Nova-3?

Ifølge dens inputskema kender Deepgram Nova-3 disse parametre: file, prompt (op til 1.000 tegn), language, temperature (0 til 1) og response_format (json, text, srt eller vtt).

Hvor hurtig er Deepgram Nova-3?

Der er endnu ikke nok målte kørsler af Deepgram Nova-3 på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.

Er Deepgram Nova-3 bedre end Incredibly Fast Whisper?

Det afhænger af opgaven. Deepgram Nova-3 (Other) og Incredibly Fast Whisper (Community) er begge modeller i kategorien Talestemme-til-tekst (STT). Sammenligningssiden viser deres priser og specifikationer side om side.

Sammenlign Deepgram Nova-3 og Incredibly Fast Whisper

Kan jeg bruge Deepgram Nova-3 lige nu?

Ikke tilgængelig i øjeblikket. Siden forbliver online; tilgængelige alternativer fra samme kategori er angivet længere nede.

Alle modeller via én API

En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.