SeamlessM4T
seamless-communicationMeta's SeamlessM4T multimodal translation model. Takes speech or text input and produces transcription or translation across about 100 languages, including speech-to-text and speech-to-speech. One model covers ASR plus cross-lingual translation without chaining separate systems.
- Price
- โ US$0.156/run
- Input โ output
- Audio โ Text
- Developer
- Community
- Updated
- 23 September 2026
Playground
Try SeamlessM4T
Input & output
This run
about US$0.156 ยท 15.6 credits
US$0.468 (46.8 credits) are reserved at the start; the actual GPU time is billed.
For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits (US$0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
Examples
- Length: 0:08
Output (JSON, shortened)
{ "text_output": "MetaAI's seamless M4T model is democratizing spoken communication across language barriers.", "audio_output": null }Output (JSON, shortened)
{ "text_output": "Eg kaller Chen Xi, pรฅ kinesisk desse to ordene betyr morgon og hรฅp โ", "audio_output": null }
About SeamlessM4T
SeamlessM4T is a model by Community in the Speech-to-text category. On Railwail, SeamlessM4T costs โ US$0.156 per run.
Pricing
| Typical run (โ 133 s on L40S) | US$0.156 per run |
|---|---|
| GPU time (L40S) | US$0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = US$0.01
Cost calculator
Price calculator
Typical according to the provider: about 133.3 s
Total
US$15.60
1,560 credits
Per run
US$0.156 ยท 15.6 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
seamless-communication- Developer
- Community
- Category
- Speech-to-text
- Input
- Audio
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- 23 September 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
input_audiorequiredURL or upload of the source audio
Type: TextDefault: โAllowed values: โtask_nameType: ChoiceDefault:ASR (Automatic Speech Recognition)Allowed values: S2ST (Speech to Speech translation), S2TT (Speech to Text translation), T2ST (Text to Speech translation), T2TT (Text to Text translation) or ASR (Automatic Speech Recognition)input_text_languageLanguage of input text when using a text-input task
Type: TextDefault: โAllowed values: โtarget_language_text_onlyTarget language for text output (e.g. English, German)
Type: TextDefault: โAllowed values: โ
Tags
- replicate
- meta
- seamless
- stt
- transcription
- translation
- multilingual
- open-weights
Use cases
Frequently asked questions
What is SeamlessM4T?
SeamlessM4T is a model by Community in the Speech-to-text category.
How much does SeamlessM4T cost on Railwail?
On Railwail, SeamlessM4T costs โ US$0.156 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.
Which settings does SeamlessM4T support?
According to its input schema, SeamlessM4T knows these parameters: input_audio, task_name (S2ST (Speech to Speech translation), S2TT (Speech to Text translation), T2ST (Text to Speech translation), T2TT (Text to Text translation) or ASR (Automatic Speech Recognition)), input_text_language and target_language_text_only.
How fast is SeamlessM4T?
There are not enough measured runs of SeamlessM4T on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is SeamlessM4T better than Incredibly Fast Whisper?
That depends on the task. SeamlessM4T (Community) and Incredibly Fast Whisper (Community) are both models in the Speech-to-text category. The comparison page shows their prices and specifications side by side.
Compare SeamlessM4T and Incredibly Fast WhisperComparable models
All in this category- Incredibly Fast WhisperCommunity
Whisper Large v3 wrapped with Hugging Face Transformers optimizations (batched inference, flash attention) for very high throughput. Transcribes hours of audio in minutes on a single GPU. Maintained by Vaibhav Srivastav. Good when you need bulk transcription fast.
- WhisperOpenAI
OpenAI's Whisper running on Replicate. General-purpose speech recognition trained on 680k hours of multilingual audio. Transcribes and translates 99 languages, robust to accents and background noise, and outputs plain text, segments, or word-level timestamps.
- SeamlessM4T v2 Large (Speech)Community
Meta SeamlessM4T v2 Large speech mode. Speech-to-speech, speech-to-text, and text-to-speech translation across 100+ languages in a single unified model.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.