F5-TTS

Konverzia textu na rečNedostupné
od X-LANCE (SJTU)ID modelu: f5-tts

Open-source flow-matching TTS with strong zero-shot voice cloning. Code MIT, weights CC-BY-NC.

Stav
Nedostupné
Vstup → výstup
Text → Audio
Vývojár
X-LANCE (SJTU)
Aktualizované
23. septembra 2026

F5-TTS nie je momentálne dostupný

Momentálne nedostupné: tento model bol deaktivovaný.

Podrobnosti na tejto stránke si môžete prečítať. Vyberte si jednu z dostupných alternatív nižšie a spustite porovnateľný model hneď.

Prejsť na alternatívy
01

Porovnateľné modely

Všetky v tejto kategórii
  • AudioLDM 2Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    ≈ 0,0157 USD/spustenie

  • Kokoro TTS 82MCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    ≈ 0,00030 USD/spustenie

  • OpenVoice v2Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    ≈ 0,0673 USD/spustenie

02

Playground

Vyskúšajte F5-TTS

Vstup a výstup

Momentálne nedostupné

Momentálne nedostupné: tento model bol deaktivovaný.

Playground je vypnutý. Porovnateľné modely nájdete v tej istej kategórii: Pozrieť alternatívy

Vyskúšajte F5-TTS

0 / 4 000

Text to speak in the cloned voice

Voice sample *Record up to 30 s · file up to 60 s, 10 MB

10 to 30 seconds of clear speech, one speaker, no music or background noise.

Pokročilé nastavenia (3)

Transcript of the reference clip (improves quality)

Výsledok
Generovaná reč sa objaví tu.

Tento beh

Bez ceny – momentálne nedostupné.

Nový tu?

10 bezplatných credits (0,10 USD) pri registrácii cez Google

Použiteľné 24 hodín po registrácii, až 5 behov za deň a maximálne 2 credits za beh. Ostatné spôsoby prihlásenia sa spúšťajú bez credits.

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
04

O F5-TTS

StručneK 23. septembra 2026

F5-TTS je model od X-LANCE (SJTU) v kategórii Konverzia textu na reč. F5-TTS nie je v súčasnosti dostupný na Railwail.

Pozadie

O SWivid (Shanghai Jiao Tong University et al.)

Založené 2024 · Shanghai, China

F5-TTS was released in October 2024 by the SWivid research collective, an open-source group anchored at Shanghai Jiao Tong University with contributors from the X-LANCE Lab, Microsoft Research Asia and the Chinese University of Hong Kong. Lead authors Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Haian Lin, Wendi He, Xiaofei Wang, Tomoki Toda, Kai Yu and Xie Chen released F5-TTS under MIT licence on GitHub together with model weights and a Hugging Face demo. The work followed the earlier E2-TTS (Microsoft) paper and quickly became one of the most popular open-weights TTS models of 2024-2025, frequently cited as the best fully open alternative to ElevenLabs for English and Chinese voice cloning.

Navštíviť SWivid (Shanghai Jiao Tong University et al.)

Architektúra

Flow-matching non-autoregressive TTS with Diffusion Transformer (DiT)

F5-TTS (Fairy-tale Fast Flow Matching TTS) is a non-autoregressive text-to-speech system that replaces the encoder-decoder pipeline with a single Diffusion Transformer trained with conditional flow matching. Text is first converted to character tokens and zero-padded to the target mel-spectrogram length, then concatenated with a noisy mel reference for the target voice. A DiT backbone with ConvNeXt v2 blocks predicts the velocity field that maps Gaussian noise to a clean mel-spectrogram, which is then converted to a waveform via Vocos vocoder. Training data is the 100k-hour open Emilia dataset (multilingual long-form audio scraped from podcasts and audiobooks). Because there is no autoregressive decoder, F5-TTS achieves real-time factor below 0.2 on a single A10 GPU while keeping competitive WER with VALL-E and NaturalSpeech 3. The model is famous for very high-quality zero-shot voice cloning from a single 5-15 second reference clip.

Parametre
~330M (Base) and ~1.4B (Large)
Kontext
30 tokenov

Schopnosti

  • Zero-shot voice cloning from ~10 s of reference audio with no fine-tuning
  • Non-autoregressive flow matching with real-time-factor < 0.2 on a single GPU
  • Open weights under MIT licence, multiple checkpoints (Base, Small, Multilingual)
  • English and Chinese out of the box; community fine-tunes for German, French, Japanese, Spanish
  • Up to 30 s of generated audio per inference
  • Speed control by adjusting the duration prompt
  • Best for: research, on-device TTS, voice cloning prototypes, commercial products built on open weights

Tréning a licencia

Pretrained on the 100,000-hour open Emilia dataset of multilingual long-form audio with weakly supervised transcripts. Additional community fine-tunes use LibriSpeech, AISHELL-3 and bespoke audiobook collections.

Licencia: Code and weights under MIT licence; commercial use permitted.

Bezpečnostné testy: No formal red-team report. Authors publish a model card recommending consent for voice cloning and a model-output watermark, which is optional.

Známe obmedzenia

  • No formal SSML / emotion tags
  • Quality degrades on noisy reference audio
  • Multilingual coverage outside English/Chinese depends on community checkpoints
  • 30-second hard cap per generation
  • Voice cloning quality slightly below ElevenLabs Multilingual V2 on emotional acting
05

Ceny

Momentálne nedostupné: tento model bol deaktivovaný. V súčasnosti nie je cena za tento model, preto ho nie je možné spustiť.

06

API

Zavolajte F5-TTS s vaším API kľúčom Railwail. V požiadavke použite toto ID modelu:

Žiadny overený príklad API

Verejné API odovzdáva iný formát vstupu, ako potrebuje tento model. Použite hraciu plochu vyššie.

07

Špecifikácie

ID modelu
f5-tts
Vývojár
X-LANCE (SJTU)
Vstup
Text
Výstup
Audio
Veľkosť modelu
~330M (Base) and ~1.4B (Large)
Licencia
Code and weights under MIT licence; commercial use permitted.
Katalógová položka aktualizovaná
23. septembra 2026

Vstupné parametre

Vstupy a nastavenia zo vstupnej schémy modelu. Príklad v sekcii API ukazuje, ktoré z nich API akceptuje.

  • audiopovinné

    Reference clip to clone; requires consent

    Typ: –
    Predvolené: –
    Povolené hodnoty: –
  • gen_textpovinné

    Text to speak in the cloned voice

    Typ: Text
    Predvolené: –
    Povolené hodnoty: až 4 000 znakov
  • speed
    Typ: Číslo
    Predvolené: 1
    Povolené hodnoty: 0,1 až 3
  • ref_text

    Transcript of the reference clip (improves quality)

    Typ: Text
    Predvolené: –
    Povolené hodnoty: –
  • remove_silence
    Typ: Áno/Nie
    Predvolené: true
    Povolené hodnoty: –

Značky

  • f5
  • tts
  • open-weights
  • voice-cloning
  • research
  • pricing-tbd
08

Prípady použitia

Na čo sa používa

  • Open-weights voice cloning research
  • On-device TTS for desktop and edge
  • Indie game and audiobook narration
  • Custom commercial products built on open weights
  • Academic benchmarks for flow-matching TTS
09

Často kladené otázky

Čo je F5-TTS?

F5-TTS je model od X-LANCE (SJTU) v kategórii Konverzia textu na reč. Je uvedený na Railwail, ale momentálne ho nie je možné spustiť.

Koľko stojí F5-TTS na Railwail?

F5-TTS nie je momentálne možné spustiť na Railwail, preto nie je aktuálna cena. Dostupné alternatívy s cenami sú uvedené nižšie na tejto stránke.

Ktoré nastavenia podporuje F5-TTS?

Podľa svojej vstupnej schémy F5-TTS pozná tieto parametre: audio, gen_text (až 4 000 znakov), speed (0,1 až 3), ref_text a remove_silence.

Ako rýchly je F5-TTS?

Pre F5-TTS je na Railwail zatiaľ príliš málo meraných spustení na určenie doby spustenia. Závisí to od vstupu, nastavení a zaťaženia u poskytovateľa.

Je F5-TTS lepší ako AudioLDM 2?

Závisí to od úlohy. F5-TTS (X-LANCE (SJTU)) a AudioLDM 2 (Haohe Liu) sú oba modely v kategórii Konverzia textu na reč. Stránka porovnania zobrazuje ich ceny a špecifikácie vedľa seba.

Porovnať F5-TTS a AudioLDM 2

Môžem F5-TTS používať práve teraz?

Momentálne nedostupné: tento model bol deaktivovaný. Stránka zostáva online; dostupné alternatívy z tej istej kategórie sú uvedené nižšie.

Všetky modely cez jedno API

Jeden API kľúč pre všetky modely na Railwail. Použitie sa účtuje z predplateného kreditu, 1 kredit = 0,01 USD.