Cartesia Sonic

Text-to-speechUnavailable
by OtherModel ID: cartesia-sonic

Cartesia's ultra-low-latency TTS (~90ms TTFB). State-space model with voice cloning support.

Status
Unavailable
Input โ†’ output
Text โ†’ Audio
Developer
Other
Updated
September 23, 2026

Cartesia Sonic is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

Go to alternatives
01

Comparable models

All in this category
02

Playground

Try Cartesia Sonic

Input & output

Currently unavailable

Currently unavailable.

The playground is disabled. You can find comparable models in the same category: Browse alternatives

Try Cartesia Sonic

0 / 4,000

Text to convert to speech

Advanced settings (2)
Output
The generated speech appears here.

This run

No price โ€“ currently unavailable.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.

03

About Cartesia Sonic

TL;DRAs of September 23, 2026

Cartesia Sonic is a model by Other in the Text-to-speech category. Cartesia Sonic is currently not available on Railwail.

Background

About Cartesia AI

Founded 2023 ยท San Francisco, California, USA

Cartesia AI was founded in 2023 by Karan Goel and Albert Gu, the academic team behind the influential state-space model line of research at Stanford and CMU (S4, H3, Mamba, Mamba-2). Co-founders also include Arjun Desai, Brandon Yang and Chris Re (Stanford advisor). The company set out to build voice and multimodal foundation models on a state-space backbone instead of Transformers in order to achieve sub-100 ms latency and stream-friendly inference. Cartesia raised a $27M seed in March 2024 led by Index Ventures with participation from Conviction, A* and Lightspeed, followed by a $64M Series A in October 2024 led by Kleiner Perkins at a reported $325M valuation. Sonic, the company's first product, launched in May 2024 and quickly became one of the lowest-latency commercial TTS systems on the market.

Visit Cartesia AI

Architecture

State-space model (Mamba family) text-to-speech with neural codec output

Cartesia Sonic is a streaming text-to-speech model built on the structured state-space modelling (SSM) architecture pioneered by Cartesia's founders (S4, H3, Mamba, Mamba-2). Unlike Transformer-based TTS systems that scale quadratically with sequence length, Sonic uses linear-time SSMs with selective scan, which lets the model maintain a small constant-memory recurrent state and generate audio chunks as text streams in. Cartesia reports a model first-byte latency of around 75-90 ms on their hosted API, which is faster than ElevenLabs Turbo and OpenAI TTS-1. Sonic outputs 24 kHz PCM via a neural codec decoder, supports streaming text input (so it can start speaking before the LLM finishes its sentence) and offers voice cloning from short reference samples (3-30 seconds). Sonic 2 added improved prosody, multilingual coverage (15+ languages) and reduced WER. The system is offered exclusively as a hosted API; weights are not released.

Parameters
Undisclosed
Context
24,000 tokens

Capabilities

  • Sub-100 ms first-byte latency suitable for real-time voice agents
  • State-space (Mamba-family) backbone with linear-time inference
  • Streaming text input and streaming audio output via WebSocket or gRPC
  • Instant voice cloning from short reference audio
  • Multilingual: English, Spanish, French, German, Portuguese, Mandarin, Japanese and more
  • Emotion and pace controls via inline tags
  • 24 kHz PCM, MP3 and Opus output formats
  • Best for: real-time voice agents, IVR systems, conversational AI, low-latency phone bots

Training & license

Cartesia has not disclosed the training corpus. Public statements describe a 'diverse multilingual speech dataset' with permissioned voice talent for the stock voice library.

License: Proprietary commercial API. Voice clones produced from customer audio remain customer property under the Terms of Service; commercial use is permitted on paid tiers.

Safety testing: Voice-cloning policy requires explicit consent statements; outputs include an inaudible watermark and audio fingerprinting to detect misuse.

Known limitations

  • Closed weights, hosted-only
  • Voice clone quality below ElevenLabs Multilingual V2 for nuanced emotional acting
  • Limited SSML / fine-grained prosody control
  • Hard cap of around 24,000 input characters per request
  • Mandarin and Japanese still less polished than English/Spanish
04

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

05

API

Call Cartesia Sonic with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
cartesia-sonic
Developer
Other
Input
Text
Output
Audio
Output formats
MP3, WAV, OPUS
Model size
Undisclosed
License
Proprietary commercial API. Voice clones produced from customer audio remain customer property under the Terms of Service; commercial use is permitted on paid tiers.
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • inputrequired

    Text to convert to speech

    Type: Text
    Default: โ€“
    Allowed values: up to 4,000 characters
  • speed
    Type: Number
    Default: 1
    Allowed values: 0.5 to 2
  • voice
    Type: Choice
    Default: sonic-default
    Allowed values: sonic-default, neutral-male, neutral-female, warm-female, or british-male
  • response_format
    Type: Choice
    Default: mp3
    Allowed values: mp3, wav, or opus

Tags

  • cartesia
  • tts
  • low-latency
  • voice-cloning
  • realtime
07

Use cases

What it is used for

  • Real-time voice agents and AI receptionists
  • Low-latency IVR / phone call automation
  • Live narration for video and streaming
  • Voice cloning for branded virtual assistants
  • Streaming TTS chained with LLM output
08

Frequently asked questions

What is Cartesia Sonic?

Cartesia Sonic is a model by Other in the Text-to-speech category. It is listed on Railwail but cannot be run at the moment.

How much does Cartesia Sonic cost on Railwail?

Cartesia Sonic cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

Which settings does Cartesia Sonic support?

According to its input schema, Cartesia Sonic knows these parameters: input (up to 4,000 characters), speed (0.5 to 2), voice (sonic-default, neutral-male, neutral-female, warm-female, or british-male), and response_format (mp3, wav, or opus).

How fast is Cartesia Sonic?

There are not enough measured runs of Cartesia Sonic on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Cartesia Sonic better than AudioLDM 2?

That depends on the task. Cartesia Sonic (Other) and AudioLDM 2 (Haohe Liu) are both models in the Text-to-speech category. The comparison page shows their prices and specifications side by side.

Compare Cartesia Sonic and AudioLDM 2

Can I use Cartesia Sonic right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.