Chatterbox

Text-to-speechAvailable
by Resemble AIModel ID: chatterbox

Resemble AI's open Chatterbox TTS. Zero-shot voice cloning from a short audio prompt with an exaggeration control for emotion intensity, plus CFG weight to balance pacing and fidelity.

Price
$0.030/1k chars
Input → output
Text → Audio
Developer
Resemble AI
Updated
September 23, 2026
01

Clone a voice

Record or upload a short sample, type a text and hear it in that voice. Chatterbox runs in your account like any other run, and you see the price first.
  1. 1

    Your voice

    Voice sample *Record up to 30 s · file up to 60 s, 10 MB

    10 to 30 seconds of clear speech, one speaker, no music or background noise.

  2. 2

    Text

    Example text, change it as you like.103 / 4,000
  3. 3

    Listen

    $0.0031 per run

02

Playground

Try Chatterbox

Input & output

$0.030/1k chars
Try Chatterbox

0 / 4,000

Text to speak

Voice sampleRecord up to 30 s · file up to 60 s, 10 MB

10 to 30 seconds of clear speech, one speaker, no music or background noise.

Advanced settings (4)

0 = random

Output
The generated speech appears here.

This run

$0.0001 · 0.01 credits

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 1,000 runs of this model.

03

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • Prompt

    We're excited to introduce Chatterbox, our first production-grade open source TTS model. Licensed under MIT, Chatterbox has been benchmarked against leading closed-source systems like ElevenLabs, and is consistently preferred in side-by-side evaluations. Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. It's also the first open source TTS model to support emotion exaggeration control, a powerful feature that makes your voices stand out. Try it now on our Hugging Face Gradio app. If you like the model but need to scale or finetune it for higher accuracy, check out our competitively priced TTS service (link). It delivers reliable performance with ultra-low latency of sub 200ms—ideal for production use in agents, applications, or interactive media.
  • Prompt

    Now let's make my mum's favourite. So three mars bars into the pan. Then we add the tuna and just stir for a bit, just let the chocolate and fish infuse. A sprinkle of olive oil and some tomato ketchup. Now smell that. Oh boy this is going to be incredible.
04

About Chatterbox

TL;DRAs of September 23, 2026

Chatterbox is a model by Resemble AI in the Text-to-speech category. On Railwail, Chatterbox costs $0.030 per 1,000 characters.

Chatterbox is Resemble AI's production-grade open TTS. It clones a voice from a brief audio_prompt and offers an exaggeration slider that scales emotional intensity, which is unusual for open models. cfg_weight and temperature tune delivery. Hosted by resemble-ai on Replicate.
05

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
1,000 characters$0.030 per 1,000 characters
  • 1 credit = $0.01

Cost calculator

Price calculator

/ run

Total

$3.00

300 credits

Per run

$0.03 · 3 credits

Fixed price per run, known before the run starts.

06

API

Call Chatterbox with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

07

Specifications

Model ID
chatterbox
Developer
Resemble AI
Input
Text
Output
Audio
Billing
Fixed price, known before the run
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • promptrequired

    Text to speak

    Type: Text
    Default: –
    Allowed values: up to 4,000 characters
  • seed

    0 = random

    Type: Integer
    Default: 0
    Allowed values: –
  • audio

    Optional reference clip (5-60 s) to clone the voice from; requires consent

    Type: –
    Default: –
    Allowed values: –
  • cfg_weight
    Type: Number
    Default: 0.5
    Allowed values: 0.2 to 1
  • temperature
    Type: Number
    Default: 0.8
    Allowed values: 0.05 to 5
  • exaggeration
    Type: Number
    Default: 0.5
    Allowed values: 0.25 to 2

Tags

  • replicate
  • resemble-ai
  • tts
  • voice-cloning
  • expressive
08

Use cases

09

Frequently asked questions

What is Chatterbox?

Chatterbox is a model by Resemble AI in the Text-to-speech category.

How much does Chatterbox cost on Railwail?

On Railwail, Chatterbox costs $0.030 per 1,000 characters. The price is known before the run starts. Usage is paid from prepaid credits; 1 credit equals $0.01.

Which settings does Chatterbox support?

According to its input schema, Chatterbox knows these parameters: prompt (up to 4,000 characters), seed, audio, cfg_weight (0.2 to 1), temperature (0.05 to 5), and exaggeration (0.25 to 2).

How fast is Chatterbox?

There are not enough measured runs of Chatterbox on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Chatterbox better than AudioLDM 2?

That depends on the task. Chatterbox (Resemble AI) and AudioLDM 2 (Haohe Liu) are both models in the Text-to-speech category. The comparison page shows their prices and specifications side by side.

Compare Chatterbox and AudioLDM 2
10

Comparable models

All in this category

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.