StyleTTS 2

Text-to-speechAvailable
by CommunityModel ID: styletts-2

Style-based TTS using diffusion and adversarial training. Human-level naturalness in zero-shot voice synthesis from a 3-5s reference clip.

Price
โ‰ˆ $0.00040/run
Input โ†’ output
Text โ†’ Audio
Developer
Community
Updated
September 23, 2026
01

Playground

Try StyleTTS 2

No input form

โ‰ˆ $0.00040/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    StyleTTS 2 is a text-to-speech model that leverages style diffusion and adversarial training with large speech language models to achieve human-level text-to-speech synthesis.
  • Prompt

    If the supply of fruit is greater than the family needs, it may be made a source of income by sending the fresh fruit to the market if there is one near enough, or by preserving, canning, and making jelly for sale. To make such an enterprise a success the fruit and work must be first class. There is magic in the word 'Homemade,' when the product appeals to the eye and the palate; but many careless and incompetent people have found to their sorrow that this word has not magic enough to float inferior goods on the market. As a rule large canning and preserving establishments are clean and have the best appliances, and they employ chemists and skilled labor. The home product must be very good to compete with the attractive goods that are sent out from such establishments. Yet for first-class homemade products there is a market in all large cities. All first-class grocers have customers who purchase such goods.
03

About StyleTTS 2

TL;DRAs of September 23, 2026

StyleTTS 2 is a model by Community in the Text-to-speech category. On Railwail, StyleTTS 2 costs โ‰ˆ $0.00040 per run.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 1 s on T4)$0.00040 per run
GPU time (T4)$0.00027 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 1.4 s

Total

$0.04

4 credits

Per run

$0.0004 ยท 0.04 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call StyleTTS 2 with your Railwail API key. Use this model ID in the request:

No verified API example

The inputs of this model are not documented yet.

06

Specifications

Model ID
styletts-2
Developer
Community
Input
Text
Output
Audio
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Tags

  • styletts
  • tts
  • voice-cloning
  • diffusion
  • open-source
07

Use cases

08

Frequently asked questions

What is StyleTTS 2?

StyleTTS 2 is a model by Community in the Text-to-speech category.

How much does StyleTTS 2 cost on Railwail?

On Railwail, StyleTTS 2 costs โ‰ˆ $0.00040 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

How fast is StyleTTS 2?

There are not enough measured runs of StyleTTS 2 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is StyleTTS 2 better than AudioLDM 2?

That depends on the task. StyleTTS 2 (Community) and AudioLDM 2 (Haohe Liu) are both models in the Text-to-speech category. The comparison page shows their prices and specifications side by side.

Compare StyleTTS 2 and AudioLDM 2
09

Comparable models

All in this category

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.