F5-TTS

Text-to-speechUnavailable
by X-LANCE (SJTU)Model ID: f5-tts

Open-source flow-matching TTS with strong zero-shot voice cloning. Code MIT, weights CC-BY-NC.

Status
Unavailable
Input β†’ output
Text β†’ Audio
Developer
X-LANCE (SJTU)
Updated
September 23, 2026

F5-TTS is currently unavailable

Currently unavailable: this model has been deactivated.

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
  • AudioLDM 2Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    β‰ˆ $0.0157/run

  • Kokoro TTS 82MCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    β‰ˆ $0.00030/run

  • OpenVoice v2Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    β‰ˆ $0.0673/run

02

Playground

Try F5-TTS

Input & output

Currently unavailable

Currently unavailable: this model has been deactivated.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try F5-TTS

0 / 4,000

Text to speak in the cloned voice

Voice sample *Record up to 30 s Β· file up to 60 s, 10 MB

10 to 30 seconds of clear speech, one speaker, no music or background noise.

Advanced settings (3)

Transcript of the reference clip (improves quality)

Output
The generated speech appears here.

This run

No price – currently unavailable.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
04

About F5-TTS

TL;DRAs of September 23, 2026

F5-TTS is a model by X-LANCE (SJTU) in the Text-to-speech category. F5-TTS is currently not available on Railwail.

Background

About SWivid (Shanghai Jiao Tong University et al.)

Founded 2024 Β· Shanghai, China

F5-TTS was released in October 2024 by the SWivid research collective, an open-source group anchored at Shanghai Jiao Tong University with contributors from the X-LANCE Lab, Microsoft Research Asia and the Chinese University of Hong Kong. Lead authors Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Haian Lin, Wendi He, Xiaofei Wang, Tomoki Toda, Kai Yu and Xie Chen released F5-TTS under MIT licence on GitHub together with model weights and a Hugging Face demo. The work followed the earlier E2-TTS (Microsoft) paper and quickly became one of the most popular open-weights TTS models of 2024-2025, frequently cited as the best fully open alternative to ElevenLabs for English and Chinese voice cloning.

Visit SWivid (Shanghai Jiao Tong University et al.)

Architecture

Flow-matching non-autoregressive TTS with Diffusion Transformer (DiT)

F5-TTS (Fairy-tale Fast Flow Matching TTS) is a non-autoregressive text-to-speech system that replaces the encoder-decoder pipeline with a single Diffusion Transformer trained with conditional flow matching. Text is first converted to character tokens and zero-padded to the target mel-spectrogram length, then concatenated with a noisy mel reference for the target voice. A DiT backbone with ConvNeXt v2 blocks predicts the velocity field that maps Gaussian noise to a clean mel-spectrogram, which is then converted to a waveform via Vocos vocoder. Training data is the 100k-hour open Emilia dataset (multilingual long-form audio scraped from podcasts and audiobooks). Because there is no autoregressive decoder, F5-TTS achieves real-time factor below 0.2 on a single A10 GPU while keeping competitive WER with VALL-E and NaturalSpeech 3. The model is famous for very high-quality zero-shot voice cloning from a single 5-15 second reference clip.

Parameters
~330M (Base) and ~1.4B (Large)
Context
30 tokens

Capabilities

  • Zero-shot voice cloning from ~10 s of reference audio with no fine-tuning
  • Non-autoregressive flow matching with real-time-factor < 0.2 on a single GPU
  • Open weights under MIT licence, multiple checkpoints (Base, Small, Multilingual)
  • English and Chinese out of the box; community fine-tunes for German, French, Japanese, Spanish
  • Up to 30 s of generated audio per inference
  • Speed control by adjusting the duration prompt
  • Best for: research, on-device TTS, voice cloning prototypes, commercial products built on open weights

Training & license

Pretrained on the 100,000-hour open Emilia dataset of multilingual long-form audio with weakly supervised transcripts. Additional community fine-tunes use LibriSpeech, AISHELL-3 and bespoke audiobook collections.

License: Code and weights under MIT licence; commercial use permitted.

Safety testing: No formal red-team report. Authors publish a model card recommending consent for voice cloning and a model-output watermark, which is optional.

Known limitations

  • No formal SSML / emotion tags
  • Quality degrades on noisy reference audio
  • Multilingual coverage outside English/Chinese depends on community checkpoints
  • 30-second hard cap per generation
  • Voice cloning quality slightly below ElevenLabs Multilingual V2 on emotional acting
05

Pricing

Currently unavailable: this model has been deactivated. There is no price for this model at the moment, so it cannot be run.

06

API

Call F5-TTS with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

07

Specifications

Model ID
f5-tts
Developer
X-LANCE (SJTU)
Input
Text
Output
Audio
Model size
~330M (Base) and ~1.4B (Large)
License
Code and weights under MIT licence; commercial use permitted.
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • audiorequired

    Reference clip to clone; requires consent

    Type: –
    Default: –
    Allowed values: –
  • gen_textrequired

    Text to speak in the cloned voice

    Type: Text
    Default: –
    Allowed values: up to 4,000 characters
  • speed
    Type: Number
    Default: 1
    Allowed values: 0.1 to 3
  • ref_text

    Transcript of the reference clip (improves quality)

    Type: Text
    Default: –
    Allowed values: –
  • remove_silence
    Type: Yes/no
    Default: true
    Allowed values: –

Tags

  • f5
  • tts
  • open-weights
  • voice-cloning
  • research
  • pricing-tbd
08

Use cases

What it is used for

  • Open-weights voice cloning research
  • On-device TTS for desktop and edge
  • Indie game and audiobook narration
  • Custom commercial products built on open weights
  • Academic benchmarks for flow-matching TTS
09

Frequently asked questions

What is F5-TTS?

F5-TTS is a model by X-LANCE (SJTU) in the Text-to-speech category. It is listed on Railwail but cannot be run at the moment.

How much does F5-TTS cost on Railwail?

F5-TTS cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

Which settings does F5-TTS support?

According to its input schema, F5-TTS knows these parameters: audio, gen_text (up to 4,000 characters), speed (0.1 to 3), ref_text, and remove_silence.

How fast is F5-TTS?

There are not enough measured runs of F5-TTS on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is F5-TTS better than AudioLDM 2?

That depends on the task. F5-TTS (X-LANCE (SJTU)) and AudioLDM 2 (Haohe Liu) are both models in the Text-to-speech category. The comparison page shows their prices and specifications side by side.

Compare F5-TTS and AudioLDM 2

Can I use F5-TTS right now?

Currently unavailable: this model has been deactivated. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.