Deepgram Nova-3

Spracherkennung (STT)Nicht verfügbar
von OtherModell-ID: nova-3

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

Status
Nicht verfügbar
Eingabe → Ausgabe
Audio → Text
Entwickler
Other
Aktualisiert
23. September 2026

Deepgram Nova-3 ist derzeit nicht verfügbar

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

Zu den Alternativen
01

Vergleichbare Modelle

Alle dieser Kategorie
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

02

Playground

Deepgram Nova-3 ausprobieren

Eingabe & Ergebnis

Derzeit nicht verfügbar

Derzeit nicht verfügbar.

Der Playground ist deaktiviert. Vergleichbare Modelle findest du in derselben Kategorie: Alternativen ansehen

Deepgram Nova-3 ausprobieren

0 / 1 000

URL or upload path to audio file

Erweiterte Einstellungen (2)
Ergebnis
Das Transkript erscheint hier.

Dieser Lauf

Kein Preis – derzeit nicht verfügbar.

Neu hier?

10 Gratis-Credits ($ 0,10) bei Anmeldung mit Google

Nutzbar 24 Stunden nach der Anmeldung, bis zu 5 Läufe pro Tag und höchstens 2 Credits je Lauf. Andere Anmeldearten starten ohne Guthaben.

03

Über Deepgram Nova-3

Kurz gesagtStand: 23. September 2026

Deepgram Nova-3 ist ein Modell von Other aus der Kategorie Spracherkennung (STT). Über Railwail ist Deepgram Nova-3 derzeit nicht verfügbar.

Hintergrund

Über Deepgram

Gegründet 2015 · San Francisco, California, USA

Deepgram was founded in 2015 by Scott Stephenson (CEO) and Noah Shutty after they finished PhDs in dark-matter physics at the University of Michigan and built a Hadoop pipeline to search audio archives. The company pivoted to commercial speech recognition in 2018 and became one of the first speech vendors to ship a fully end-to-end deep-learning ASR pipeline rather than the classical Kaldi stack. Deepgram has raised over $86M in venture funding from Tiger Global, Y Combinator, Wing, Madrona and NVIDIA's NVentures, with a 2024 Series C at a reported $1B+ valuation. The Nova model family launched in 2023 (Nova-1), followed by Nova-2 (2024) and Nova-3 (January 2025), positioning Deepgram as the fastest commercial ASR provider with sub-300 ms streaming and self-hostable deployments.

Deepgram besuchen

Architektur

End-to-end encoder-decoder speech-to-text with self-supervised pretraining

Deepgram Nova-3 is an end-to-end deep-learning automatic speech recognition model trained on what Deepgram describes as the largest commercial training corpus among production ASR providers, including 50,000+ hours of curated multi-domain audio plus self-supervised pretraining on hundreds of thousands of hours of unlabelled speech. Nova-3 introduced a unified multilingual model that natively handles English, Spanish, French, German, Italian, Portuguese, Hindi, Japanese, Mandarin, Korean and 10+ more, with on-the-fly code-switching inside a single audio file. New capabilities include keyword prompting (custom vocabulary at request time without retraining), Self-Hosted edition for HIPAA/SOC2 deployments, and improved noise robustness on call-centre and phone audio. The streaming variant runs at sub-300 ms latency over WebSocket, while the pre-recorded variant supports up to 4 hours per file. Outputs include diarisation, punctuation, smart formatting, language detection and topic/intent metadata.

Parameter
Undisclosed (multi-billion)
Kontext
14 400 Token

Funktionen

  • Real-time streaming ASR at sub-300 ms latency
  • Multilingual single model with code-switching across 36+ languages
  • Keyword prompting for runtime custom vocabulary
  • Speaker diarisation, punctuation and smart formatting
  • Phone-call optimised (8 kHz, low-bandwidth, noisy)
  • Self-Hosted deployment for HIPAA / SOC2 / on-premises
  • Up to 4 hours of audio per pre-recorded request
  • Best for: call centres, real-time meeting transcription, voice agents, compliance recording

Training & Lizenz

Self-supervised pretraining on hundreds of thousands of hours of unlabelled audio plus supervised training on 50,000+ hours of curated, labelled multilingual speech. Data is sourced from licensed corpora, customer opt-in audio and publicly available datasets.

Lizenz: Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.

Sicherheitstests: SOC 2 Type II and HIPAA compliant; no model-output watermarking required for transcripts; bias testing reported in the Nova-3 launch blog.

Bekannte Einschränkungen

  • Word Error Rate worse than Whisper Large v3 on some long-form audiobook tasks
  • Streaming diarisation can mislabel speakers in fast cross-talk
  • Self-Hosted edition has high hardware requirements
  • Code-switching quality varies between language pairs
  • Closed weights for hosted API
04

Preise

Derzeit nicht verfügbar. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

05

API

Rufe Deepgram Nova-3 mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Kein geprüftes API-Beispiel

Die öffentliche API übergibt ein anderes Eingabeformat, als dieses Modell braucht. Nutze den Playground oben.

06

Spezifikationen

Modell-ID
nova-3
Entwickler
Other
Eingabe
Audio
Ausgabe
Text
Ausgabeformate
JSON, SRT, VTT
Modellgröße
Undisclosed (multi-billion)
Lizenz
Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.
Katalogeintrag aktualisiert
23. September 2026

Eingabeparameter

Eingaben und Einstellungen laut Eingabeschema des Modells. Welche davon die API annimmt, zeigt das Beispiel im Abschnitt API.

  • filePflicht

    URL or upload path to audio file

    Typ: Text
    Standard: –
    Erlaubte Werte: –
  • prompt

    Optional context to guide transcription

    Typ: Text
    Standard: –
    Erlaubte Werte: bis 1 000 Zeichen
  • language

    Optional ISO-639-1 language code (e.g. en, de, fr)

    Typ: Text
    Standard: –
    Erlaubte Werte: –
  • temperature
    Typ: Zahl
    Standard: 0
    Erlaubte Werte: 0 bis 1
  • response_format
    Typ: Auswahl
    Standard: json
    Erlaubte Werte: json, text, srt oder vtt

Schlagwörter

  • deepgram
  • stt
  • transcription
  • realtime
  • multilingual
  • per-minute
07

Einsatzgebiete

Wofür es genutzt wird

  • Real-time call-centre transcription and analytics
  • Voice agent transcript layer
  • Live closed captioning
  • Compliance recording for finance and healthcare
  • Meeting transcription with diarisation
08

Häufige Fragen

Was ist Deepgram Nova-3?

Deepgram Nova-3 ist ein Modell von Other aus der Kategorie Spracherkennung (STT). Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet Deepgram Nova-3 bei Railwail?

Deepgram Nova-3 lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Welche Einstellungen unterstützt Deepgram Nova-3?

Laut Eingabeschema kennt Deepgram Nova-3 diese Parameter: file, prompt (bis 1 000 Zeichen), language, temperature (0 bis 1) und response_format (json, text, srt oder vtt).

Wie schnell ist Deepgram Nova-3?

Für Deepgram Nova-3 gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Ist Deepgram Nova-3 besser als Incredibly Fast Whisper?

Das hängt von der Aufgabe ab. Deepgram Nova-3 (Other) und Incredibly Fast Whisper (Community) sind beide Modelle aus der Kategorie Spracherkennung (STT). Die Vergleichsseite zeigt Preise und Spezifikationen nebeneinander.

Deepgram Nova-3 und Incredibly Fast Whisper vergleichen

Kann ich Deepgram Nova-3 gerade nutzen?

Derzeit nicht verfügbar. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0,01.