ElevenLabs Scribe v1

Reconnaissance vocaleNon disponible
par ElevenLabsID du modèle: scribe-v1

ElevenLabs' STT. 99 languages, word-level timestamps, speaker diarization, audio-event tagging.

Statut
Non disponible
Entrée → Sortie
Audio → Texte
Développeur
ElevenLabs
Mis à jour
25 juin 2026

ElevenLabs Scribe v1 n'est actuellement pas disponible

Actuellement indisponible : ce modèle a été désactivé.

Vous pouvez toujours consulter les détails sur cette page. Choisissez l'une des alternatives disponibles ci-dessous pour exécuter immédiatement un modèle comparable.

Voir les alternatives
01

Modèles comparables

Tous dans cette catégorie
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ 0,0056 $US/exécution

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ 0,0034 $US/exécution

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ 0,156 $US/exécution

02

Playground

Essayer ElevenLabs Scribe v1

Pas de formulaire d'entrée

Actuellement indisponible

Actuellement indisponible : ce modèle a été désactivé.

Le playground est désactivé. Vous pouvez trouver des modèles comparables dans la même catégorie : Parcourir les alternatives

03

À propos de ElevenLabs Scribe v1

RésuméAu 25 juin 2026

ElevenLabs Scribe v1 est un modèle de ElevenLabs dans la catégorie Reconnaissance vocale. ElevenLabs Scribe v1 n'est actuellement pas disponible sur Railwail.

Arrière-plan

À propos de ElevenLabs

Fondée 2022 · London, UK / New York, USA

ElevenLabs was founded in 2022 by Piotr Dabkowski and Mati Staniszewski, two Polish technologists who had worked at Google and Palantir. The company became famous for high-quality multilingual text-to-speech and AI dubbing, and in February 2025 expanded into the inverse problem with Scribe v1, the company's first dedicated automatic speech recognition model. Scribe was developed in part to power ElevenLabs' Dubbing Studio (transcribe source audio, translate, then re-synthesise in the target language), and is offered as a standalone API to enterprise customers who want a single vendor for the full STT-to-TTS pipeline. ElevenLabs has raised over $280M to date, with a Series C in January 2025 at a $3.3B valuation.

Visiter ElevenLabs

Architecture

Proprietary encoder-decoder speech-to-text Transformer

ElevenLabs Scribe v1 is a hosted automatic speech recognition model launched in February 2025. ElevenLabs has not published a technical report, but the launch blog describes a Transformer encoder-decoder ASR architecture trained on a large multilingual speech corpus covering 99 languages, with particular emphasis on accuracy in long-tail languages where Whisper Large v3 underperforms. Scribe outperformed Whisper Large v3 and Deepgram Nova-2 in the company's published FLEURS and Common Voice evaluations across many language pairs, and ranked first overall in a head-to-head benchmark on Hindi, Mandarin, German and Italian. The model supports speaker diarisation up to 32 speakers, word-level timestamps with sub-100 ms precision, character-level confidence scores, automatic non-speech event detection ([applause], [laughter], [music]) and audio-event classification. Maximum file size is 1 GB and maximum audio length is 2 hours per request.

Paramètres
Undisclosed
Contexte
7 200 tokens

Capacités

  • 99-language multilingual ASR including many low-resource languages
  • Speaker diarisation up to 32 speakers
  • Word-level timestamps with sub-100 ms precision
  • Non-speech event detection ([applause], [laughter], [music])
  • Character-level confidence scores
  • Up to 2 hours per request, 1 GB file limit
  • Direct integration with ElevenLabs Dubbing Studio (STT to translate to TTS)
  • Best for: dubbing pipelines, multilingual transcription, podcast indexing, media analytics

Entraînement et licence

Not disclosed. ElevenLabs reports training on a 'large multilingual corpus' with curation for long-tail languages; data is described as a mix of licensed and crowd-sourced opt-in audio.

Licence: Proprietary commercial API. Commercial use permitted on paid plans.

Tests de sécurité: Same KYC and identity-verification practices as the rest of the ElevenLabs platform; no formal red-team report for ASR.

Limitations connues

  • No streaming mode at launch (file-based only)
  • Hard cap of 2 hours per request
  • Pricing per minute higher than Deepgram Nova-3 for English
  • Closed weights, hosted only
  • Diarisation accuracy degrades in noisy cross-talk
04

Tarification

Actuellement indisponible : ce modèle a été désactivé. Il n'y a actuellement pas de prix pour ce modèle, il ne peut donc pas être exécuté.

05

API

Appelez ElevenLabs Scribe v1 avec votre clé API Railwail. Utilisez cet ID de modèle dans la requête :

Aucun exemple API vérifié

L'API publique transmet un format d'entrée différent de celui dont ce modèle a besoin. Utilisez le playground ci-dessus.

06

Spécifications

ID du modèle
scribe-v1
Développeur
ElevenLabs
Entrée
Audio
Sortie
Texte
Taille du modèle
Undisclosed
Licence
Proprietary commercial API. Commercial use permitted on paid plans.
Entrée du catalogue mise à jour
25 juin 2026

Étiquettes

  • elevenlabs
  • scribe
  • stt
  • transcription
  • diarization
  • per-minute
07

Cas d'usage

À quoi ça sert

  • AI dubbing pipelines (STT to translate to TTS)
  • Multilingual podcast and media transcription
  • Search and indexing of audio archives
  • Meeting transcription with speaker diarisation
  • Accessibility captioning for video
08

Questions fréquemment posées

Qu'est-ce que ElevenLabs Scribe v1 ?

ElevenLabs Scribe v1 est un modèle de ElevenLabs dans la catégorie Reconnaissance vocale. Il est répertorié sur Railwail mais ne peut pas être exécuté pour le moment.

Combien coûte ElevenLabs Scribe v1 sur Railwail ?

ElevenLabs Scribe v1 ne peut pas être exécuté sur Railwail pour le moment, il n'y a donc pas de prix actuel. Les alternatives disponibles avec leurs prix sont listées plus bas sur cette page.

Quelle est la vitesse de ElevenLabs Scribe v1 ?

Il n'y a pas encore assez d'exécutions mesurées de ElevenLabs Scribe v1 sur Railwail pour indiquer un temps d'exécution. Cela dépend de l'entrée, des paramètres et de la charge chez le fournisseur.

ElevenLabs Scribe v1 est-il meilleur que Incredibly Fast Whisper ?

Cela dépend de la tâche. ElevenLabs Scribe v1 (ElevenLabs) et Incredibly Fast Whisper (Community) sont tous deux des modèles de la catégorie Reconnaissance vocale. La page de comparaison affiche leurs prix et spécifications côte à côte.

Comparer ElevenLabs Scribe v1 et Incredibly Fast Whisper

Puis-je utiliser ElevenLabs Scribe v1 maintenant ?

Actuellement indisponible : ce modèle a été désactivé. La page reste en ligne ; les alternatives disponibles de la même catégorie sont listées plus bas.

Tous les modèles via une seule API

Une clé API pour tous les modèles sur Railwail. L'utilisation est facturée à partir de crédits prépayés, 1 crédit = 0,01 $US.