F5-TTS

Převod textu na řečNedostupné
od X-LANCE (SJTU)ID modelu: f5-tts

Open-source flow-matching TTS with strong zero-shot voice cloning. Code MIT, weights CC-BY-NC.

Stav
Nedostupné
Vstup → výstup
Text → Audio
Vývojář
X-LANCE (SJTU)
Aktualizováno
23. září 2026

F5-TTS není momentálně dostupný

Momentálně nedostupné: tento model byl deaktivován.

Podrobnosti na této stránce si můžete přečíst. Vyberte jednu z dostupných alternativ níže a spusťte porovnatelný model hned.

Na alternativy
01

Srovnatelné modely

Všechny v této kategorii
  • AudioLDM 2Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    ≈ 0,0157 US$/spuštění

  • Kokoro TTS 82MCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    ≈ 0,00030 US$/spuštění

  • OpenVoice v2Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    ≈ 0,0673 US$/spuštění

02

Playground

Vyzkoušet F5-TTS

Vstup a výstup

Momentálně nedostupné

Momentálně nedostupné: tento model byl deaktivován.

Playground je zakázán. Srovnatelné modely najdete ve stejné kategorii: Zobrazit alternativy

Vyzkoušet F5-TTS

0 / 4 000

Text to speak in the cloned voice

Voice sample *Record up to 30 s · file up to 60 s, 10 MB

10 to 30 seconds of clear speech, one speaker, no music or background noise.

Pokročilá nastavení (3)

Transcript of the reference clip (improves quality)

Výsledek
Vygenerovaná řeč se zobrazí zde.

Tento běh

Bez ceny – momentálně nedostupné.

Nový zde?

10 bezplatných credits (0,10 US$) při registraci přes Google

Použitelné 24 hodin po registraci, až 5 běhů za den a maximálně 2 credits za běh. Ostatní způsoby přihlášení začínají bez credits.

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
04

O F5-TTS

StručněStav: 23. září 2026

F5-TTS je model od X-LANCE (SJTU) v kategorii Převod textu na řeč. F5-TTS není na Railwail v současné době dostupný.

Pozadí

O SWivid (Shanghai Jiao Tong University et al.)

Založeno 2024 · Shanghai, China

F5-TTS was released in October 2024 by the SWivid research collective, an open-source group anchored at Shanghai Jiao Tong University with contributors from the X-LANCE Lab, Microsoft Research Asia and the Chinese University of Hong Kong. Lead authors Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Haian Lin, Wendi He, Xiaofei Wang, Tomoki Toda, Kai Yu and Xie Chen released F5-TTS under MIT licence on GitHub together with model weights and a Hugging Face demo. The work followed the earlier E2-TTS (Microsoft) paper and quickly became one of the most popular open-weights TTS models of 2024-2025, frequently cited as the best fully open alternative to ElevenLabs for English and Chinese voice cloning.

Navštívit SWivid (Shanghai Jiao Tong University et al.)

Architektura

Flow-matching non-autoregressive TTS with Diffusion Transformer (DiT)

F5-TTS (Fairy-tale Fast Flow Matching TTS) is a non-autoregressive text-to-speech system that replaces the encoder-decoder pipeline with a single Diffusion Transformer trained with conditional flow matching. Text is first converted to character tokens and zero-padded to the target mel-spectrogram length, then concatenated with a noisy mel reference for the target voice. A DiT backbone with ConvNeXt v2 blocks predicts the velocity field that maps Gaussian noise to a clean mel-spectrogram, which is then converted to a waveform via Vocos vocoder. Training data is the 100k-hour open Emilia dataset (multilingual long-form audio scraped from podcasts and audiobooks). Because there is no autoregressive decoder, F5-TTS achieves real-time factor below 0.2 on a single A10 GPU while keeping competitive WER with VALL-E and NaturalSpeech 3. The model is famous for very high-quality zero-shot voice cloning from a single 5-15 second reference clip.

Parametry
~330M (Base) and ~1.4B (Large)
Kontext
30 tokenů

Schopnosti

  • Zero-shot voice cloning from ~10 s of reference audio with no fine-tuning
  • Non-autoregressive flow matching with real-time-factor < 0.2 on a single GPU
  • Open weights under MIT licence, multiple checkpoints (Base, Small, Multilingual)
  • English and Chinese out of the box; community fine-tunes for German, French, Japanese, Spanish
  • Up to 30 s of generated audio per inference
  • Speed control by adjusting the duration prompt
  • Best for: research, on-device TTS, voice cloning prototypes, commercial products built on open weights

Trénování a licence

Pretrained on the 100,000-hour open Emilia dataset of multilingual long-form audio with weakly supervised transcripts. Additional community fine-tunes use LibriSpeech, AISHELL-3 and bespoke audiobook collections.

Licence: Code and weights under MIT licence; commercial use permitted.

Bezpečnostní testy: No formal red-team report. Authors publish a model card recommending consent for voice cloning and a model-output watermark, which is optional.

Známá omezení

  • No formal SSML / emotion tags
  • Quality degrades on noisy reference audio
  • Multilingual coverage outside English/Chinese depends on community checkpoints
  • 30-second hard cap per generation
  • Voice cloning quality slightly below ElevenLabs Multilingual V2 on emotional acting
05

Ceny

Momentálně nedostupné: tento model byl deaktivován. Pro tento model momentálně není cena, takže jej nelze spustit.

06

API

Volejte F5-TTS s vaším API klíčem Railwail. V požadavku použijte toto ID modelu:

Žádný ověřený příklad API

Veřejné API předává jiný formát vstupu, než tento model potřebuje. Použijte playground výše.

07

Specifikace

ID modelu
f5-tts
Vývojář
X-LANCE (SJTU)
Vstup
Text
Výstup
Audio
Velikost modelu
~330M (Base) and ~1.4B (Large)
Licence
Code and weights under MIT licence; commercial use permitted.
Záznam v katalogu aktualizován
23. září 2026

Vstupní parametry

Vstupy a nastavení ze vstupního schématu modelu. Příklad v sekci API ukazuje, které z nich API přijímá.

  • audiopovinné

    Reference clip to clone; requires consent

    Typ: –
    Výchozí: –
    Povolené hodnoty: –
  • gen_textpovinné

    Text to speak in the cloned voice

    Typ: Text
    Výchozí: –
    Povolené hodnoty: až 4 000 znaků
  • speed
    Typ: Číslo
    Výchozí: 1
    Povolené hodnoty: 0,1 až 3
  • ref_text

    Transcript of the reference clip (improves quality)

    Typ: Text
    Výchozí: –
    Povolené hodnoty: –
  • remove_silence
    Typ: Ano/Ne
    Výchozí: true
    Povolené hodnoty: –

Štítky

  • f5
  • tts
  • open-weights
  • voice-cloning
  • research
  • pricing-tbd
08

Případy použití

K čemu se používá

  • Open-weights voice cloning research
  • On-device TTS for desktop and edge
  • Indie game and audiobook narration
  • Custom commercial products built on open weights
  • Academic benchmarks for flow-matching TTS
09

Často kladené otázky

Co je F5-TTS?

F5-TTS je model od X-LANCE (SJTU) v kategorii Převod textu na řeč. Je uveden na Railwail, ale momentálně jej nelze spustit.

Kolik stojí F5-TTS na Railwail?

F5-TTS se momentálně na Railwail nedá spustit, takže není aktuální cena. Dostupné alternativy s cenami jsou uvedeny dále na této stránce.

Jaká nastavení F5-TTS podporuje?

Podle schématu vstupu F5-TTS zná tyto parametry: audio, gen_text (až 4 000 znaků), speed (0,1 až 3), ref_text a remove_silence.

Jak rychlý je F5-TTS?

Pro F5-TTS je na Railwail zatím příliš málo naměřených spuštění na uvedení doby běhu. Závisí na vstupu, nastavení a zátěži u poskytovatele.

Je F5-TTS lepší než AudioLDM 2?

Záleží na úkolu. F5-TTS (X-LANCE (SJTU)) a AudioLDM 2 (Haohe Liu) jsou oba modely v kategorii Převod textu na řeč. Stránka porovnání zobrazuje jejich ceny a specifikace vedle sebe.

Porovnat F5-TTS a AudioLDM 2

Mohu F5-TTS používat hned teď?

Momentálně nedostupné: tento model byl deaktivován. Stránka zůstává online; dostupné alternativy ze stejné kategorie jsou uvedeny dále.

Všechny modely přes jednu API

Jeden API klíč pro všechny modely na Railwail. Použití se účtuje z předplacených kreditů, 1 kredit = 0,01 US$.