SeamlessM4T

Riconoscimento vocale (STT)Disponibile
di CommunityID modello: seamless-communication

Meta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.

Prezzo
≈ 0,156 USD/esecuzione
Input → output
Audio → Testo
Sviluppatore
Community
Aggiornato
23 settembre 2026
01

Playground

Prova SeamlessM4T

Input e output

≈ 0,156 USD/esecuzione
Prova SeamlessM4T

URL or upload of the source audio

Impostazioni avanzate (3)

Language of input text when using a text-input task

Target language for text output (e.g. English, German)

Risultato
La trascrizione appare qui.

Questa esecuzione

circa 0,156 USD · 15,6 crediti

0,468 USD (46,8 crediti) sono riservati all'inizio; il tempo GPU effettivo viene fatturato.

Per account senza acquisti precedenti: le esecuzioni oltre 2 crediti richiedono una ricarica.

Nuovo qui?

10 crediti gratuiti (0,10 USD) quando ti iscrivi con Google

Utilizzabile 24 ore dopo l'iscrizione, fino a 5 esecuzioni al giorno e al massimo 2 crediti per esecuzione. Altri metodi di accesso iniziano senza crediti.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Output (JSON, shortened)

    {
      "text_output": "MetaAI's seamless M4T model is democratizing spoken communication across language barriers.",
      "audio_output": null
    }
  • Output (JSON, shortened)

    {
      "text_output": "Eg kaller Chen Xi, på kinesisk desse to ordene betyr morgon og håp ⁇",
      "audio_output": null
    }
03

Informazioni su SeamlessM4T

RiassuntoA partire da 23 settembre 2026

SeamlessM4T è un modello di Community nella categoria Riconoscimento vocale (STT). Su Railwail, SeamlessM4T costa ≈ 0,156 USD per esecuzione.

04

Prezzi

Prezzi in dollari USA. L'utilizzo viene addebitato dai crediti prepagati.
Esecuzione tipica (≈ 133 s su L40S)0,156 USD per esecuzione
Tempo GPU (L40S)0,00117 USD per secondo GPU
  • Fatturato in base al tempo GPU effettivamente utilizzato. All'avvio, 3× il prezzo tipico viene riservato dal tuo saldo e regolato successivamente.
  • 1 credito = 0,01 USD

Calcolatore di costi

Calcolatore prezzi

s

Tipico secondo il provider: circa 133,3 s

Totale

15,60 USD

1560 crediti

Per esecuzione

0,156 USD · 15,6 crediti

Fatturato in base al tempo GPU effettivo; questo è una stima.

05

API

Chiama SeamlessM4T con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Nessun esempio API verificato

L'API pubblica passa un formato di input diverso da quello richiesto da questo modello. Usa il playground sopra.

06

Specifiche

ID modello
seamless-communication
Sviluppatore
Community
Input
Audio
Output
Testo
Fatturazione
In base all'utilizzo (token o tempo GPU)
Voce di catalogo aggiornata
23 settembre 2026

Parametri di input

Input e impostazioni dallo schema di input del modello. L'esempio nella sezione API mostra quali di essi l'API accetta.

  • input_audioObbligatorio

    URL or upload of the source audio

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • task_name
    Tipo: Scelta
    Predefinito: ASR (Automatic Speech Recognition)
    Valori consentiti: S2ST (Speech to Speech translation), S2TT (Speech to Text translation), T2ST (Text to Speech translation), T2TT (Text to Text translation) o ASR (Automatic Speech Recognition)
  • input_text_language

    Language of input text when using a text-input task

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –
  • target_language_text_only

    Target language for text output (e.g. English, German)

    Tipo: Testo
    Predefinito: –
    Valori consentiti: –

Etichette

  • replicate
  • meta
  • seamless
  • stt
  • transcription
  • translation
  • multilingual
  • open-weights
07

Casi d'uso

08

Domande frequenti

Cos'è SeamlessM4T?

SeamlessM4T è un modello di Community nella categoria Riconoscimento vocale (STT).

Quanto costa SeamlessM4T su Railwail?

Su Railwail, SeamlessM4T costa ≈ 0,156 USD per esecuzione. Ti viene addebitato ciò che ogni richiesta utilizza effettivamente. L'utilizzo viene pagato con crediti prepagati; 1 credito equivale a 0,01 USD.

Quali impostazioni supporta SeamlessM4T?

Secondo il suo schema di input, SeamlessM4T conosce questi parametri: input_audio, task_name (S2ST (Speech to Speech translation), S2TT (Speech to Text translation), T2ST (Text to Speech translation), T2TT (Text to Text translation) o ASR (Automatic Speech Recognition)), input_text_language e target_language_text_only.

Quanto è veloce SeamlessM4T?

Non ci sono ancora abbastanza esecuzioni misurate di SeamlessM4T su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

SeamlessM4T è migliore di Incredibly Fast Whisper?

Dipende dall'attività. SeamlessM4T (Community) e Incredibly Fast Whisper (Community) sono entrambi modelli nella categoria Riconoscimento vocale (STT). La pagina di confronto mostra i loro prezzi e le specifiche affiancati.

Confronta SeamlessM4T e Incredibly Fast Whisper
09

Modelli comparabili

Tutti in questa categoria
  • Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.

    ≈ 0,0056 USD/esecuzione

    96 % più economico per unità

    Confronta SeamlessM4T e Incredibly Fast Whisper
  • WhisperOpenAI

    OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.

    ≈ 0,0034 USD/esecuzione

    98 % più economico per unità

    Confronta SeamlessM4T e Whisper
  • Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.

    ≈ 0,0012 USD/esecuzione

    99 % più economico per unità

    Confronta SeamlessM4T e SeamlessM4T v2 Large (Speech)

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.