StyleTTS 2

Text-to-speechAvailable
by CommunityModel ID: styletts-2

Style-based TTS using diffusion and adversarial training. Human-level naturalness in zero-shot voice synthesis from a 3-5s reference clip.

Price
โ‰ˆ US$0.00040/run
Input โ†’ output
Text โ†’ Audio
Developer
Community
Updated
23 September 2026
01

Playground

Try StyleTTS 2

No input form

โ‰ˆ US$0.00040/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    StyleTTS 2 is a text-to-speech model that leverages style diffusion and adversarial training with large speech language models to achieve human-level text-to-speech synthesis.
  • Prompt

    If the supply of fruit is greater than the family needs, it may be made a source of income by sending the fresh fruit to the market if there is one near enough, or by preserving, canning, and making jelly for sale. To make such an enterprise a success the fruit and work must be first class. There is magic in the word 'Homemade,' when the product appeals to the eye and the palate; but many careless and incompetent people have found to their sorrow that this word has not magic enough to float inferior goods on the market. As a rule large canning and preserving establishments are clean and have the best appliances, and they employ chemists and skilled labor. The home product must be very good to compete with the attractive goods that are sent out from such establishments. Yet for first-class homemade products there is a market in all large cities. All first-class grocers have customers who purchase such goods.
03

About StyleTTS 2

TL;DRAs of 23 September 2026

StyleTTS 2 is a model by Community in the Text-to-speech category. On Railwail, StyleTTS 2 costs โ‰ˆ US$0.00040 per run.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 1 s on T4)US$0.00040 per run
GPU time (T4)US$0.00027 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = US$0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 1.4 s

Total

US$0.04

4 credits

Per run

US$0.0004 ยท 0.04 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call StyleTTS 2 with your Railwail API key. Use this model ID in the request:

No verified API example

The inputs of this model are not documented yet.

06

Specifications

Model ID
styletts-2
Developer
Community
Input
Text
Output
Audio
Billing
By usage (tokens or GPU time)
Catalog entry updated
23 September 2026

Tags

  • styletts
  • tts
  • voice-cloning
  • diffusion
  • open-source
07

Use cases

08

Frequently asked questions

What is StyleTTS 2?

StyleTTS 2 is a model by Community in the Text-to-speech category.

How much does StyleTTS 2 cost on Railwail?

On Railwail, StyleTTS 2 costs โ‰ˆ US$0.00040 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.

How fast is StyleTTS 2?

There are not enough measured runs of StyleTTS 2 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is StyleTTS 2 better than AudioLDM 2?

That depends on the task. StyleTTS 2 (Community) and AudioLDM 2 (Haohe Liu) are both models in the Text-to-speech category. The comparison page shows their prices and specifications side by side.

Compare StyleTTS 2 and AudioLDM 2
09

Comparable models

All in this category
  • AudioLDM 2Haohe Liu

    Latent-diffusion model for general-purpose text-to-audio. Generates speech, music, and sound effects with a unified prior.

    โ‰ˆ US$0.0157/run

    3,825 % more expensive per unit

    Compare StyleTTS 2 vs. AudioLDM 2
  • Kokoro TTS 82MCommunity

    Open-weights 82M-parameter TTS. Punches above its size class on naturalness benchmarks at a fraction of the inference cost of larger models.

    โ‰ˆ US$0.00030/run

    25 % cheaper per unit

    Compare StyleTTS 2 vs. Kokoro TTS 82M
  • OpenVoice v2Community

    MyShell OpenVoice v2. Multilingual zero-shot voice cloning with accurate tone-color reproduction and style/emotion control.

    โ‰ˆ US$0.0673/run

    16,725 % more expensive per unit

    Compare StyleTTS 2 vs. OpenVoice v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.