ElevenLabs Scribe v1

Conversão de fala em textoIndisponível
por ElevenLabsID do modelo: scribe-v1

ElevenLabs' STT. 99 languages, word-level timestamps, speaker diarization, audio-event tagging.

Status
Indisponível
Entrada → Saída
Áudio → Texto
Desenvolvedor
ElevenLabs
Atualizado
25 de junho de 2026

ElevenLabs Scribe v1 não está disponível no momento

Atualmente indisponível: este modelo foi desativado.

Você ainda pode ler os detalhes nesta página. Escolha uma das alternativas disponíveis abaixo para executar um modelo comparável imediatamente.

Ir para alternativas
01

Modelos comparáveis

Todos nesta categoria
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ US$ 0,0056/execução

  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ US$ 0,0034/execução

  • SeamlessM4TCommunity

    Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

    ≈ US$ 0,156/execução

02

Playground

Experimentar ElevenLabs Scribe v1

Sem formulário de entrada

Indisponível no momento

Atualmente indisponível: este modelo foi desativado.

O playground está desativado. Encontre modelos comparáveis na mesma categoria: Ver alternativas

03

Sobre ElevenLabs Scribe v1

ResumoA partir de 25 de junho de 2026

ElevenLabs Scribe v1 é um modelo de ElevenLabs na categoria Conversão de fala em texto. ElevenLabs Scribe v1 não está disponível no Railwail no momento.

Fundo

Sobre ElevenLabs

Fundado em 2022 · London, UK / New York, USA

ElevenLabs was founded in 2022 by Piotr Dabkowski and Mati Staniszewski, two Polish technologists who had worked at Google and Palantir. The company became famous for high-quality multilingual text-to-speech and AI dubbing, and in February 2025 expanded into the inverse problem with Scribe v1, the company's first dedicated automatic speech recognition model. Scribe was developed in part to power ElevenLabs' Dubbing Studio (transcribe source audio, translate, then re-synthesise in the target language), and is offered as a standalone API to enterprise customers who want a single vendor for the full STT-to-TTS pipeline. ElevenLabs has raised over $280M to date, with a Series C in January 2025 at a $3.3B valuation.

Visite ElevenLabs

Arquitetura

Proprietary encoder-decoder speech-to-text Transformer

ElevenLabs Scribe v1 is a hosted automatic speech recognition model launched in February 2025. ElevenLabs has not published a technical report, but the launch blog describes a Transformer encoder-decoder ASR architecture trained on a large multilingual speech corpus covering 99 languages, with particular emphasis on accuracy in long-tail languages where Whisper Large v3 underperforms. Scribe outperformed Whisper Large v3 and Deepgram Nova-2 in the company's published FLEURS and Common Voice evaluations across many language pairs, and ranked first overall in a head-to-head benchmark on Hindi, Mandarin, German and Italian. The model supports speaker diarisation up to 32 speakers, word-level timestamps with sub-100 ms precision, character-level confidence scores, automatic non-speech event detection ([applause], [laughter], [music]) and audio-event classification. Maximum file size is 1 GB and maximum audio length is 2 hours per request.

Parâmetros
Undisclosed
Contexto
7.200 tokens

Capacidades

  • 99-language multilingual ASR including many low-resource languages
  • Speaker diarisation up to 32 speakers
  • Word-level timestamps with sub-100 ms precision
  • Non-speech event detection ([applause], [laughter], [music])
  • Character-level confidence scores
  • Up to 2 hours per request, 1 GB file limit
  • Direct integration with ElevenLabs Dubbing Studio (STT to translate to TTS)
  • Best for: dubbing pipelines, multilingual transcription, podcast indexing, media analytics

Treinamento & licença

Not disclosed. ElevenLabs reports training on a 'large multilingual corpus' with curation for long-tail languages; data is described as a mix of licensed and crowd-sourced opt-in audio.

Licença: Proprietary commercial API. Commercial use permitted on paid plans.

Testes de segurança: Same KYC and identity-verification practices as the rest of the ElevenLabs platform; no formal red-team report for ASR.

Limitações conhecidas

  • No streaming mode at launch (file-based only)
  • Hard cap of 2 hours per request
  • Pricing per minute higher than Deepgram Nova-3 for English
  • Closed weights, hosted only
  • Diarisation accuracy degrades in noisy cross-talk
04

Preços

Atualmente indisponível: este modelo foi desativado. Não há preço para este modelo no momento, portanto não pode ser executado.

05

API

Chame ElevenLabs Scribe v1 com sua chave de API Railwail. Use este ID de modelo na solicitação:

Nenhum exemplo de API verificado

A API pública passa um formato de entrada diferente do que este modelo precisa. Use o playground acima.

06

Especificações

ID do modelo
scribe-v1
Desenvolvedor
ElevenLabs
Entrada
Áudio
Saída
Texto
Tamanho do modelo
Undisclosed
Licença
Proprietary commercial API. Commercial use permitted on paid plans.
Entrada do catálogo atualizada
25 de junho de 2026

Etiquetas

  • elevenlabs
  • scribe
  • stt
  • transcription
  • diarization
  • per-minute
07

Casos de uso

Para que é utilizado

  • AI dubbing pipelines (STT to translate to TTS)
  • Multilingual podcast and media transcription
  • Search and indexing of audio archives
  • Meeting transcription with speaker diarisation
  • Accessibility captioning for video
08

Perguntas frequentes

O que é ElevenLabs Scribe v1?

ElevenLabs Scribe v1 é um modelo de ElevenLabs na categoria Conversão de fala em texto. Está listado no Railwail, mas não pode ser executado no momento.

Quanto custa ElevenLabs Scribe v1 no Railwail?

ElevenLabs Scribe v1 não pode ser executado no Railwail no momento, portanto não há preço atual. Alternativas disponíveis com preços estão listadas mais abaixo nesta página.

Qual é a velocidade de ElevenLabs Scribe v1?

Ainda não há execuções medidas suficientes de ElevenLabs Scribe v1 no Railwail para indicar um tempo de execução. Depende da entrada, das configurações e da carga no provedor.

ElevenLabs Scribe v1 é melhor que Incredibly Fast Whisper?

Depende da tarefa. ElevenLabs Scribe v1 (ElevenLabs) e Incredibly Fast Whisper (Community) são ambos modelos na categoria Conversão de fala em texto. A página de comparação mostra seus preços e especificações lado a lado.

Comparar ElevenLabs Scribe v1 e Incredibly Fast Whisper

Posso usar ElevenLabs Scribe v1 agora?

Atualmente indisponível: este modelo foi desativado. A página permanece online; alternativas disponíveis da mesma categoria estão listadas mais abaixo.

Todos os modelos através de uma API

Uma chave API para todos os modelos no Railwail. O uso é cobrado a partir de créditos pré-pagos, 1 crédito = US$ 0,01.