Deepgram Nova-3

Reconnaissance vocaleNon disponible
par OtherID du modèle: nova-3

Deepgram's flagship STT. First to offer realtime multilingual transcription with self-serve customization.

Statut
Non disponible
Entrée → Sortie
Audio → Texte
Développeur
Other
Mis à jour
23 septembre 2026

Deepgram Nova-3 n'est actuellement pas disponible

Vous pouvez toujours consulter les détails sur cette page. Choisissez l'une des alternatives disponibles ci-dessous pour exécuter immédiatement un modèle comparable.

Voir les alternatives
01

Modèles comparables

Tous dans cette catégorie
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ 0,0056 $US/exécution

    Comparer Deepgram Nova-3 et Incredibly Fast Whisper
  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ 0,0034 $US/exécution

    Comparer Deepgram Nova-3 et Whisper
  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ 0,156 $US/exécution

    Comparer Deepgram Nova-3 et SeamlessM4T
02

Playground

Essayer Deepgram Nova-3

Entrée et résultat

Actuellement indisponible

Actuellement indisponible.

Le playground est désactivé. Vous pouvez trouver des modèles comparables dans la même catégorie : Parcourir les alternatives

Essayer Deepgram Nova-3

0 / 1 000

URL or upload path to audio file

Paramètres avancés (2)
Résultat
La transcription apparaît ici.

Cette exécution

Pas de prix – actuellement indisponible.

Nouveau par ici ?

10 crédits gratuits (0,10 $US) lors de l'inscription avec Google

Utilisable 24 heures après l'inscription, jusqu'à 5 exécutions par jour et au maximum 2 crédits par exécution. Les autres méthodes de connexion commencent sans crédits.

03

À propos de Deepgram Nova-3

RésuméAu 23 septembre 2026

Deepgram Nova-3 est un modèle de Other dans la catégorie Reconnaissance vocale. Deepgram Nova-3 n'est actuellement pas disponible sur Railwail.

Arrière-plan

À propos de Deepgram

Fondée 2015 · San Francisco, California, USA

Deepgram was founded in 2015 by Scott Stephenson (CEO) and Noah Shutty after they finished PhDs in dark-matter physics at the University of Michigan and built a Hadoop pipeline to search audio archives. The company pivoted to commercial speech recognition in 2018 and became one of the first speech vendors to ship a fully end-to-end deep-learning ASR pipeline rather than the classical Kaldi stack. Deepgram has raised over $86M in venture funding from Tiger Global, Y Combinator, Wing, Madrona and NVIDIA's NVentures, with a 2024 Series C at a reported $1B+ valuation. The Nova model family launched in 2023 (Nova-1), followed by Nova-2 (2024) and Nova-3 (January 2025), positioning Deepgram as the fastest commercial ASR provider with sub-300 ms streaming and self-hostable deployments.

Visiter Deepgram

Architecture

End-to-end encoder-decoder speech-to-text with self-supervised pretraining

Deepgram Nova-3 is an end-to-end deep-learning automatic speech recognition model trained on what Deepgram describes as the largest commercial training corpus among production ASR providers, including 50,000+ hours of curated multi-domain audio plus self-supervised pretraining on hundreds of thousands of hours of unlabelled speech. Nova-3 introduced a unified multilingual model that natively handles English, Spanish, French, German, Italian, Portuguese, Hindi, Japanese, Mandarin, Korean and 10+ more, with on-the-fly code-switching inside a single audio file. New capabilities include keyword prompting (custom vocabulary at request time without retraining), Self-Hosted edition for HIPAA/SOC2 deployments, and improved noise robustness on call-centre and phone audio. The streaming variant runs at sub-300 ms latency over WebSocket, while the pre-recorded variant supports up to 4 hours per file. Outputs include diarisation, punctuation, smart formatting, language detection and topic/intent metadata.

Paramètres
Undisclosed (multi-billion)
Contexte
14 400 tokens

Capacités

  • Real-time streaming ASR at sub-300 ms latency
  • Multilingual single model with code-switching across 36+ languages
  • Keyword prompting for runtime custom vocabulary
  • Speaker diarisation, punctuation and smart formatting
  • Phone-call optimised (8 kHz, low-bandwidth, noisy)
  • Self-Hosted deployment for HIPAA / SOC2 / on-premises
  • Up to 4 hours of audio per pre-recorded request
  • Best for: call centres, real-time meeting transcription, voice agents, compliance recording

Entraînement et licence

Self-supervised pretraining on hundreds of thousands of hours of unlabelled audio plus supervised training on 50,000+ hours of curated, labelled multilingual speech. Data is sourced from licensed corpora, customer opt-in audio and publicly available datasets.

Licence: Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.

Tests de sécurité: SOC 2 Type II and HIPAA compliant; no model-output watermarking required for transcripts; bias testing reported in the Nova-3 launch blog.

Limitations connues

  • Word Error Rate worse than Whisper Large v3 on some long-form audiobook tasks
  • Streaming diarisation can mislabel speakers in fast cross-talk
  • Self-Hosted edition has high hardware requirements
  • Code-switching quality varies between language pairs
  • Closed weights for hosted API
04

Tarification

Actuellement indisponible. Il n'y a actuellement pas de prix pour ce modèle, il ne peut donc pas être exécuté.

05

API

Appelez Deepgram Nova-3 avec votre clé API Railwail. Utilisez cet ID de modèle dans la requête :

Aucun exemple API vérifié

L'API publique transmet un format d'entrée différent de celui dont ce modèle a besoin. Utilisez le playground ci-dessus.

06

Spécifications

ID du modèle
nova-3
Développeur
Other
Entrée
Audio
Sortie
Texte
Formats de sortie
JSON, SRT, VTT
Taille du modèle
Undisclosed (multi-billion)
Licence
Proprietary commercial API and Self-Hosted licence. Commercial use is permitted; Self-Hosted requires a separate enterprise agreement.
Entrée du catalogue mise à jour
23 septembre 2026

Paramètres d'entrée

Entrées et paramètres du schéma d'entrée du modèle. L'exemple dans la section API montre lesquels l'API accepte.

  • fileObligatoire

    URL or upload path to audio file

    Type: Texte
    Défaut: –
    Valeurs autorisées: –
  • prompt

    Optional context to guide transcription

    Type: Texte
    Défaut: –
    Valeurs autorisées: Jusqu'à 1 000 caractères
  • language

    Optional ISO-639-1 language code (e.g. en, de, fr)

    Type: Texte
    Défaut: –
    Valeurs autorisées: –
  • temperature
    Type: Nombre
    Défaut: 0
    Valeurs autorisées: 0 à 1
  • response_format
    Type: Choix
    Défaut: json
    Valeurs autorisées: json, text, srt ou vtt

Étiquettes

  • deepgram
  • stt
  • transcription
  • realtime
  • multilingual
  • per-minute
07

Cas d'usage

À quoi ça sert

  • Real-time call-centre transcription and analytics
  • Voice agent transcript layer
  • Live closed captioning
  • Compliance recording for finance and healthcare
  • Meeting transcription with diarisation
08

Questions fréquemment posées

Qu'est-ce que Deepgram Nova-3 ?

Deepgram Nova-3 est un modèle de Other dans la catégorie Reconnaissance vocale. Il est répertorié sur Railwail mais ne peut pas être exécuté pour le moment.

Combien coûte Deepgram Nova-3 sur Railwail ?

Deepgram Nova-3 ne peut pas être exécuté sur Railwail pour le moment, il n'y a donc pas de prix actuel. Les alternatives disponibles avec leurs prix sont listées plus bas sur cette page.

Quels paramètres Deepgram Nova-3 supporte-t-il ?

Selon son schéma d'entrée, Deepgram Nova-3 connaît ces paramètres : file, prompt (Jusqu'à 1 000 caractères), language, temperature (0 à 1) et response_format (json, text, srt ou vtt).

Quelle est la vitesse de Deepgram Nova-3 ?

Il n'y a pas encore assez d'exécutions mesurées de Deepgram Nova-3 sur Railwail pour indiquer un temps d'exécution. Cela dépend de l'entrée, des paramètres et de la charge chez le fournisseur.

Deepgram Nova-3 est-il meilleur que Incredibly Fast Whisper ?

Cela dépend de la tâche. Deepgram Nova-3 (Other) et Incredibly Fast Whisper (Community) sont tous deux des modèles de la catégorie Reconnaissance vocale. La page de comparaison affiche leurs prix et spécifications côte à côte.

Comparer Deepgram Nova-3 et Incredibly Fast Whisper

Puis-je utiliser Deepgram Nova-3 maintenant ?

Actuellement indisponible. La page reste en ligne ; les alternatives disponibles de la même catégorie sont listées plus bas.

Tous les modèles via une seule API

Une clé API pour tous les modèles sur Railwail. L'utilisation est facturée à partir de crédits prépayés, 1 crédit = 0,01 $US.