ElevenLabs Scribe v1

Reconocimiento de vozNo disponible
de ElevenLabsID del modelo: scribe-v1

ElevenLabs' STT. 99 languages, word-level timestamps, speaker diarization, audio-event tagging.

Estado
No disponible
Entrada → Salida
Audio → Texto
Desarrollador
ElevenLabs
Actualizado
25 de junio de 2026

ElevenLabs Scribe v1 no está disponible en este momento

Actualmente no disponible: este modelo ha sido desactivado.

Puedes seguir leyendo los detalles en esta página. Elige una de las alternativas disponibles abajo para ejecutar un modelo comparable de inmediato.

Ir a alternativas
01

Modelos comparables

Todos en esta categoría
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ USD 0.0056/ejecución

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ USD 0.0034/ejecución

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ USD 0.156/ejecución

02

Playground

Probar ElevenLabs Scribe v1

Sin formulario de entrada

No disponible actualmente

Actualmente no disponible: este modelo ha sido desactivado.

El área de pruebas está desactivada. Encuentra modelos comparables en la misma categoría: Explorar alternativas

03

Acerca de ElevenLabs Scribe v1

ResumenA partir de 25 de junio de 2026

ElevenLabs Scribe v1 es un modelo de ElevenLabs en la categoría Reconocimiento de voz. ElevenLabs Scribe v1 no está disponible en Railwail en este momento.

Fondo

Acerca de ElevenLabs

Fundado en 2022 · London, UK / New York, USA

ElevenLabs was founded in 2022 by Piotr Dabkowski and Mati Staniszewski, two Polish technologists who had worked at Google and Palantir. The company became famous for high-quality multilingual text-to-speech and AI dubbing, and in February 2025 expanded into the inverse problem with Scribe v1, the company's first dedicated automatic speech recognition model. Scribe was developed in part to power ElevenLabs' Dubbing Studio (transcribe source audio, translate, then re-synthesise in the target language), and is offered as a standalone API to enterprise customers who want a single vendor for the full STT-to-TTS pipeline. ElevenLabs has raised over $280M to date, with a Series C in January 2025 at a $3.3B valuation.

Visitar ElevenLabs

Arquitectura

Proprietary encoder-decoder speech-to-text Transformer

ElevenLabs Scribe v1 is a hosted automatic speech recognition model launched in February 2025. ElevenLabs has not published a technical report, but the launch blog describes a Transformer encoder-decoder ASR architecture trained on a large multilingual speech corpus covering 99 languages, with particular emphasis on accuracy in long-tail languages where Whisper Large v3 underperforms. Scribe outperformed Whisper Large v3 and Deepgram Nova-2 in the company's published FLEURS and Common Voice evaluations across many language pairs, and ranked first overall in a head-to-head benchmark on Hindi, Mandarin, German and Italian. The model supports speaker diarisation up to 32 speakers, word-level timestamps with sub-100 ms precision, character-level confidence scores, automatic non-speech event detection ([applause], [laughter], [music]) and audio-event classification. Maximum file size is 1 GB and maximum audio length is 2 hours per request.

Parámetros
Undisclosed
Contexto
7,200 tokens

Capacidades

  • 99-language multilingual ASR including many low-resource languages
  • Speaker diarisation up to 32 speakers
  • Word-level timestamps with sub-100 ms precision
  • Non-speech event detection ([applause], [laughter], [music])
  • Character-level confidence scores
  • Up to 2 hours per request, 1 GB file limit
  • Direct integration with ElevenLabs Dubbing Studio (STT to translate to TTS)
  • Best for: dubbing pipelines, multilingual transcription, podcast indexing, media analytics

Entrenamiento y licencia

Not disclosed. ElevenLabs reports training on a 'large multilingual corpus' with curation for long-tail languages; data is described as a mix of licensed and crowd-sourced opt-in audio.

Licencia: Proprietary commercial API. Commercial use permitted on paid plans.

Pruebas de seguridad: Same KYC and identity-verification practices as the rest of the ElevenLabs platform; no formal red-team report for ASR.

Limitaciones conocidas

  • No streaming mode at launch (file-based only)
  • Hard cap of 2 hours per request
  • Pricing per minute higher than Deepgram Nova-3 for English
  • Closed weights, hosted only
  • Diarisation accuracy degrades in noisy cross-talk
04

Precios

Actualmente no disponible: este modelo ha sido desactivado. No hay precio para este modelo en este momento, por lo que no se puede ejecutar.

05

API

Llama a ElevenLabs Scribe v1 con tu clave de API de Railwail. Usa este ID de modelo en la solicitud:

Sin ejemplo de API verificado

La API pública pasa un formato de entrada diferente al que necesita este modelo. Usa el playground anterior.

06

Especificaciones

ID del modelo
scribe-v1
Desarrollador
ElevenLabs
Entrada
Audio
Salida
Texto
Tamaño del modelo
Undisclosed
Licencia
Proprietary commercial API. Commercial use permitted on paid plans.
Entrada del catálogo actualizada
25 de junio de 2026

Etiquetas

  • elevenlabs
  • scribe
  • stt
  • transcription
  • diarization
  • per-minute
07

Casos de uso

Para qué se utiliza

  • AI dubbing pipelines (STT to translate to TTS)
  • Multilingual podcast and media transcription
  • Search and indexing of audio archives
  • Meeting transcription with speaker diarisation
  • Accessibility captioning for video
08

Preguntas frecuentes

¿Qué es ElevenLabs Scribe v1?

ElevenLabs Scribe v1 es un modelo de ElevenLabs en la categoría Reconocimiento de voz. Está listado en Railwail pero no se puede ejecutar en este momento.

¿Cuánto cuesta ElevenLabs Scribe v1 en Railwail?

ElevenLabs Scribe v1 no se puede ejecutar en Railwail en este momento, por lo que no hay precio actual. Las alternativas disponibles con precios se enumeran más abajo en esta página.

¿Qué tan rápido es ElevenLabs Scribe v1?

Aún no hay suficientes ejecuciones medidas de ElevenLabs Scribe v1 en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.

¿Es ElevenLabs Scribe v1 mejor que Incredibly Fast Whisper?

Eso depende de la tarea. ElevenLabs Scribe v1 (ElevenLabs) y Incredibly Fast Whisper (Community) son ambos modelos en la categoría Reconocimiento de voz. La página de comparación muestra sus precios y especificaciones lado a lado.

Comparar ElevenLabs Scribe v1 y Incredibly Fast Whisper

¿Puedo usar ElevenLabs Scribe v1 ahora mismo?

Actualmente no disponible: este modelo ha sido desactivado. La página permanece en línea; las alternativas disponibles de la misma categoría se enumeran más abajo.

Todos los modelos a través de una API

Una clave API para todos los modelos en Railwail. El uso se cobra con créditos prepagados, 1 crédito = USD 0.01.