F5-TTS

Teksti puheeksiEi saatavilla
kehittäjä: X-LANCE (SJTU)Mallin tunnus: f5-tts

Open-source flow-matching TTS with strong zero-shot voice cloning. Code MIT, weights CC-BY-NC.

Tila
Ei saatavilla
Syöte → tulos
Teksti → Audio
Kehittäjä
X-LANCE (SJTU)
Päivitetty
23. syyskuuta 2026

F5-TTS ei ole tällä hetkellä saatavilla

Ei tällä hetkellä saatavilla: tämä malli on poistettu käytöstä.

Voit silti lukea tiedot tältä sivulta. Valitse yksi alla olevista saatavilla olevista vaihtoehdoista suorittaaksesi vertailukelpoisen mallin heti.

Siirry vaihtoehtoihin
01

Vertailukelpoiset mallit

Kaikki tässä kategoriassa
  • AudioLDM 2Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    ≈ 0,0157 $/suoritus

  • Kokoro TTS 82MCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    ≈ 0,00030 $/suoritus

  • OpenVoice v2Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    ≈ 0,0673 $/suoritus

02

Leikkikenttä

Kokeile F5-TTS

Syöte ja tulos

Ei tällä hetkellä saatavilla

Ei tällä hetkellä saatavilla: tämä malli on poistettu käytöstä.

Leikkikenttä on poistettu käytöstä. Vertailukelpoisia malleja löydät samasta kategoriasta: Selaa vaihtoehtoja

Kokeile F5-TTS

0 / 4 000

Text to speak in the cloned voice

Voice sample *Record up to 30 s · file up to 60 s, 10 MB

10 to 30 seconds of clear speech, one speaker, no music or background noise.

Lisäasetukset (3)

Transcript of the reference clip (improves quality)

Tulos
Luotu puhe ilmestyy tähän.

Tämä suoritus

Ei hintaa – ei tällä hetkellä saatavilla.

Uusi täällä?

10 ilmaista creditiä (0,10 $) kun rekisteröidyt Googlella

Käytettävissä 24 tuntia rekisteröinnin jälkeen, enintään 5 suoritusta päivässä ja enintään 2 creditiä suoritusta kohti. Muut kirjautumismenetelmät alkavat ilman creditejä.

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
04

Tietoja: F5-TTS

Lyhyesti23. syyskuuta 2026 alkaen

F5-TTS on X-LANCE (SJTU)-kehittäjän malli kategoriasta Teksti puheeksi. F5-TTS ei ole tällä hetkellä saatavilla Railwailissa.

Tausta

Tietoja: SWivid (Shanghai Jiao Tong University et al.)

Perustettu 2024 · Shanghai, China

F5-TTS was released in October 2024 by the SWivid research collective, an open-source group anchored at Shanghai Jiao Tong University with contributors from the X-LANCE Lab, Microsoft Research Asia and the Chinese University of Hong Kong. Lead authors Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Haian Lin, Wendi He, Xiaofei Wang, Tomoki Toda, Kai Yu and Xie Chen released F5-TTS under MIT licence on GitHub together with model weights and a Hugging Face demo. The work followed the earlier E2-TTS (Microsoft) paper and quickly became one of the most popular open-weights TTS models of 2024-2025, frequently cited as the best fully open alternative to ElevenLabs for English and Chinese voice cloning.

Vieraile sivustolla SWivid (Shanghai Jiao Tong University et al.)

Arkkitehtuuri

Flow-matching non-autoregressive TTS with Diffusion Transformer (DiT)

F5-TTS (Fairy-tale Fast Flow Matching TTS) is a non-autoregressive text-to-speech system that replaces the encoder-decoder pipeline with a single Diffusion Transformer trained with conditional flow matching. Text is first converted to character tokens and zero-padded to the target mel-spectrogram length, then concatenated with a noisy mel reference for the target voice. A DiT backbone with ConvNeXt v2 blocks predicts the velocity field that maps Gaussian noise to a clean mel-spectrogram, which is then converted to a waveform via Vocos vocoder. Training data is the 100k-hour open Emilia dataset (multilingual long-form audio scraped from podcasts and audiobooks). Because there is no autoregressive decoder, F5-TTS achieves real-time factor below 0.2 on a single A10 GPU while keeping competitive WER with VALL-E and NaturalSpeech 3. The model is famous for very high-quality zero-shot voice cloning from a single 5-15 second reference clip.

Parametrit
~330M (Base) and ~1.4B (Large)
Konteksti
30 tokenia

Ominaisuudet

  • Zero-shot voice cloning from ~10 s of reference audio with no fine-tuning
  • Non-autoregressive flow matching with real-time-factor < 0.2 on a single GPU
  • Open weights under MIT licence, multiple checkpoints (Base, Small, Multilingual)
  • English and Chinese out of the box; community fine-tunes for German, French, Japanese, Spanish
  • Up to 30 s of generated audio per inference
  • Speed control by adjusting the duration prompt
  • Best for: research, on-device TTS, voice cloning prototypes, commercial products built on open weights

Koulutus ja lisenssi

Pretrained on the 100,000-hour open Emilia dataset of multilingual long-form audio with weakly supervised transcripts. Additional community fine-tunes use LibriSpeech, AISHELL-3 and bespoke audiobook collections.

Lisenssi: Code and weights under MIT licence; commercial use permitted.

Turvallisuustestit: No formal red-team report. Authors publish a model card recommending consent for voice cloning and a model-output watermark, which is optional.

Tunnetut rajoitukset

  • No formal SSML / emotion tags
  • Quality degrades on noisy reference audio
  • Multilingual coverage outside English/Chinese depends on community checkpoints
  • 30-second hard cap per generation
  • Voice cloning quality slightly below ElevenLabs Multilingual V2 on emotional acting
05

Hinnat

Ei tällä hetkellä saatavilla: tämä malli on poistettu käytöstä. Tällä hetkellä tälle mallille ei ole hintaa, joten sitä ei voi suorittaa.

06

API

Kutsu F5-TTS Railwail-API-avaimellasi. Käytä tätä mallin tunnusta pyynnössä:

Ei vahvistettua API-esimerkkiä

Julkinen API välittää eri syötemuotoa kuin tämä malli tarvitsee. Käytä yllä olevaa leikkikenttää.

07

Tekniset tiedot

Mallin tunnus
f5-tts
Kehittäjä
X-LANCE (SJTU)
Syöte
Teksti
Tuloste
Audio
Mallin koko
~330M (Base) and ~1.4B (Large)
Lisenssi
Code and weights under MIT licence; commercial use permitted.
Luettelokirjaus päivitetty
23. syyskuuta 2026

Syöteparametrit

Mallin syötökaavion syötteet ja asetukset. API-osion esimerkki näyttää, mitkä niistä API hyväksyy.

  • audiopakollinen

    Reference clip to clone; requires consent

    Tyyppi: –
    Oletus: –
    Sallitut arvot: –
  • gen_textpakollinen

    Text to speak in the cloned voice

    Tyyppi: Teksti
    Oletus: –
    Sallitut arvot: enintään 4 000 merkkiä
  • speed
    Tyyppi: Luku
    Oletus: 1
    Sallitut arvot: 0,1–3
  • ref_text

    Transcript of the reference clip (improves quality)

    Tyyppi: Teksti
    Oletus: –
    Sallitut arvot: –
  • remove_silence
    Tyyppi: Kyllä/Ei
    Oletus: true
    Sallitut arvot: –

Tunnisteet

  • f5
  • tts
  • open-weights
  • voice-cloning
  • research
  • pricing-tbd
08

Käyttötapaukset

Mihin sitä käytetään

  • Open-weights voice cloning research
  • On-device TTS for desktop and edge
  • Indie game and audiobook narration
  • Custom commercial products built on open weights
  • Academic benchmarks for flow-matching TTS
09

Usein kysytyt kysymykset

Mikä on F5-TTS?

F5-TTS on X-LANCE (SJTU)n kehittämä malli Teksti puheeksi-kategoriassa. Se on listattu Railwailissa, mutta sitä ei voi tällä hetkellä suorittaa.

Paljonko F5-TTS maksaa Railwailissa?

F5-TTSta ei voi tällä hetkellä suorittaa Railwailissa, joten nykyistä hintaa ei ole. Saatavilla olevat vaihtoehdot hintojen kanssa on lueteltu tämän sivun alempana.

Mitä asetuksia F5-TTS tukee?

Syötteen rakenteen mukaan F5-TTS tuntee nämä parametrit: audio, gen_text (enintään 4 000 merkkiä), speed (0,1–3), ref_text ja remove_silence.

Kuinka nopea F5-TTS on?

F5-TTSlla ei ole vielä tarpeeksi mitattuja suorituksia Railwailissa suoritusajan ilmoittamiseksi. Se riippuu syötteestä, asetuksista ja palveluntarjoajan kuormituksesta.

Onko F5-TTS parempi kuin AudioLDM 2?

Se riippuu tehtävästä. F5-TTS (X-LANCE (SJTU)) ja AudioLDM 2 (Haohe Liu) ovat molemmat malleja Teksti puheeksi-kategoriassa. Vertailussa näkyvät niiden hinnat ja tekniset tiedot rinnakkain.

Vertaa F5-TTS ja AudioLDM 2

Voiko F5-TTSa käyttää juuri nyt?

Ei tällä hetkellä saatavilla: tämä malli on poistettu käytöstä. Sivu pysyy verkossa; saatavilla olevat vaihtoehdot samasta kategoriasta on lueteltu alempana.

Kaikki mallit yhden API:n kautta

Yksi API-avain kaikille Railwailin malleille. Käyttö laskutetaan prepaid-krediiteistä, 1 krediitti = 0,01 $.