F5-TTS

Sinteză vocală (TTS)Indisponibil
de X-LANCE (SJTU)ID model: f5-tts

Open-source flow-matching TTS with strong zero-shot voice cloning. Code MIT, weights CC-BY-NC.

Status
Indisponibil
Intrare → ieșire
Text → Audio
Dezvoltator
X-LANCE (SJTU)
Actualizat
23 septembrie 2026

F5-TTS nu este disponibil în acest moment

Indisponibil în prezent: acest model a fost dezactivat.

Poți citi în continuare detaliile pe această pagină. Alege una dintre alternativele disponibile de mai jos pentru a rula imediat un model comparabil.

Mergi la alternative
01
  • AudioLDM 2Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    ≈ 0,0157 USD/rulare

  • Kokoro TTS 82MCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    ≈ 0,00030 USD/rulare

  • OpenVoice v2Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    ≈ 0,0673 USD/rulare

02

Playground

Încearcă F5-TTS

Intrare & rezultat

Indisponibil în prezent

Indisponibil în prezent: acest model a fost dezactivat.

Playground-ul este dezactivat. Modele comparabile găsești în aceeași categorie: Vezi alternativele

Încearcă F5-TTS

0 / 4.000

Text to speak in the cloned voice

Voice sample *Record up to 30 s · file up to 60 s, 10 MB

10 to 30 seconds of clear speech, one speaker, no music or background noise.

Setări avansate (3)

Transcript of the reference clip (improves quality)

Rezultat
Vocea generată apare aici.

Această rulare

Fără preț – momentan indisponibil.

Nou aici?

10 credite gratuite (0,10 USD) când te înregistrezi cu Google

Utilizabil 24 ore după înregistrare, până la 5 rulări pe zi și maximum 2 credite pe rulare. Alte metode de conectare încep fără credite.

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
04

Despre F5-TTS

Pe scurtDin 23 septembrie 2026

F5-TTS este un model de X-LANCE (SJTU) din categoria Sinteză vocală (TTS). F5-TTS nu este disponibil în prezent pe Railwail.

Fundal

Despre SWivid (Shanghai Jiao Tong University et al.)

Fondat 2024 · Shanghai, China

F5-TTS was released in October 2024 by the SWivid research collective, an open-source group anchored at Shanghai Jiao Tong University with contributors from the X-LANCE Lab, Microsoft Research Asia and the Chinese University of Hong Kong. Lead authors Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Haian Lin, Wendi He, Xiaofei Wang, Tomoki Toda, Kai Yu and Xie Chen released F5-TTS under MIT licence on GitHub together with model weights and a Hugging Face demo. The work followed the earlier E2-TTS (Microsoft) paper and quickly became one of the most popular open-weights TTS models of 2024-2025, frequently cited as the best fully open alternative to ElevenLabs for English and Chinese voice cloning.

Vizitează SWivid (Shanghai Jiao Tong University et al.)

Arhitectură

Flow-matching non-autoregressive TTS with Diffusion Transformer (DiT)

F5-TTS (Fairy-tale Fast Flow Matching TTS) is a non-autoregressive text-to-speech system that replaces the encoder-decoder pipeline with a single Diffusion Transformer trained with conditional flow matching. Text is first converted to character tokens and zero-padded to the target mel-spectrogram length, then concatenated with a noisy mel reference for the target voice. A DiT backbone with ConvNeXt v2 blocks predicts the velocity field that maps Gaussian noise to a clean mel-spectrogram, which is then converted to a waveform via Vocos vocoder. Training data is the 100k-hour open Emilia dataset (multilingual long-form audio scraped from podcasts and audiobooks). Because there is no autoregressive decoder, F5-TTS achieves real-time factor below 0.2 on a single A10 GPU while keeping competitive WER with VALL-E and NaturalSpeech 3. The model is famous for very high-quality zero-shot voice cloning from a single 5-15 second reference clip.

Parametri
~330M (Base) and ~1.4B (Large)
Context
30 tokeni

Capabilități

  • Zero-shot voice cloning from ~10 s of reference audio with no fine-tuning
  • Non-autoregressive flow matching with real-time-factor < 0.2 on a single GPU
  • Open weights under MIT licence, multiple checkpoints (Base, Small, Multilingual)
  • English and Chinese out of the box; community fine-tunes for German, French, Japanese, Spanish
  • Up to 30 s of generated audio per inference
  • Speed control by adjusting the duration prompt
  • Best for: research, on-device TTS, voice cloning prototypes, commercial products built on open weights

Antrenament & licență

Pretrained on the 100,000-hour open Emilia dataset of multilingual long-form audio with weakly supervised transcripts. Additional community fine-tunes use LibriSpeech, AISHELL-3 and bespoke audiobook collections.

Licență: Code and weights under MIT licence; commercial use permitted.

Teste de siguranță: No formal red-team report. Authors publish a model card recommending consent for voice cloning and a model-output watermark, which is optional.

Limitări cunoscute

  • No formal SSML / emotion tags
  • Quality degrades on noisy reference audio
  • Multilingual coverage outside English/Chinese depends on community checkpoints
  • 30-second hard cap per generation
  • Voice cloning quality slightly below ElevenLabs Multilingual V2 on emotional acting
05

Prețuri

Indisponibil în prezent: acest model a fost dezactivat. Nu există preț pentru acest model în acest moment, deci nu poate fi executat.

06

API

Apelează F5-TTS cu cheia ta API Railwail. Folosește acest ID de model în cerere:

Niciun exemplu API verificat

API-ul public transmite un format de intrare diferit de ceea ce are nevoie acest model. Folosește playground-ul de mai sus.

07

Specificații

ID model
f5-tts
Dezvoltator
X-LANCE (SJTU)
Intrare
Text
Ieșire
Audio
Dimensiune model
~330M (Base) and ~1.4B (Large)
Licență
Code and weights under MIT licence; commercial use permitted.
Intrare catalog actualizată
23 septembrie 2026

Parametri de intrare

Intrări și setări din schema de intrare a modelului. Exemplul din secțiunea API arată care dintre ele acceptă API-ul.

  • audioobligatoriu

    Reference clip to clone; requires consent

    Tip: –
    Implicit: –
    Valori permise: –
  • gen_textobligatoriu

    Text to speak in the cloned voice

    Tip: Text
    Implicit: –
    Valori permise: până la 4.000 caractere
  • speed
    Tip: Număr
    Implicit: 1
    Valori permise: 0,1 până la 3
  • ref_text

    Transcript of the reference clip (improves quality)

    Tip: Text
    Implicit: –
    Valori permise: –
  • remove_silence
    Tip: Da/Nu
    Implicit: true
    Valori permise: –

Etichete

  • f5
  • tts
  • open-weights
  • voice-cloning
  • research
  • pricing-tbd
08

Cazuri de utilizare

Pentru ce se folosește

  • Open-weights voice cloning research
  • On-device TTS for desktop and edge
  • Indie game and audiobook narration
  • Custom commercial products built on open weights
  • Academic benchmarks for flow-matching TTS
09

Întrebări frecvente

Ce este F5-TTS?

F5-TTS este un model de X-LANCE (SJTU) din categoria Sinteză vocală (TTS). Este listat pe Railwail, dar nu poate fi rulat în acest moment.

Cât costă F5-TTS pe Railwail?

F5-TTS nu poate fi rulat pe Railwail în acest moment, deci nu există preț curent. Alternativele disponibile cu prețuri sunt listate mai jos pe această pagină.

Ce setări acceptă F5-TTS?

Conform schemei sale de intrare, F5-TTS cunoaște acești parametri: audio, gen_text (până la 4.000 caractere), speed (0,1 până la 3), ref_text și remove_silence.

Cât de rapid este F5-TTS?

Nu sunt suficiente rulări măsurate ale F5-TTS pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.

Este F5-TTS mai bun decât AudioLDM 2?

Depinde de sarcină. F5-TTS (X-LANCE (SJTU)) și AudioLDM 2 (Haohe Liu) sunt ambele modele din categoria Sinteză vocală (TTS). Pagina de comparație arată prețurile și specificațiile lor una lângă alta.

Compară F5-TTS și AudioLDM 2

Pot folosi F5-TTS chiar acum?

Indisponibil în prezent: acest model a fost dezactivat. Pagina rămâne online; alternativele disponibile din aceeași categorie sunt listate mai jos.

Toate modelele printr-o singură API

O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.