Skip to main content

Cartesia Ships Sonic-3.6 Streaming TTS Model

Cartesia launches Sonic-3.6, a streaming text-to-speech model featuring sub-90ms latency and leading Artificial Analysis speech arenas.

AI-written
Inewgen
18 Aug 2026Source: MarkTechPost3 min read (0 views)Last updated 29 Aug 2026
Share
Cartesia Ships Sonic-3.6 Streaming TTS Model

Stock photo for illustration only, not from the actual event

Font size
  • Cartesia releases its latest streaming TTS model, Sonic-3.6, in beta and as a hosted API
  • Achieves sub-90ms TTS latency and 100ms transcript latency for the Ink-2 STT model
  • Powered by state space models instead of traditional transformer architectures
  • Priced at $49.00 per 1 million characters according to Artificial Analysis benchmarks

Artificial intelligence developer Cartesia has officially announced the launch of its newest text-to-speech model, Sonic-3.6. The model is currently available in beta and through a hosted API as a closed, commercial offering. It features no open weights or Hugging Face repository, requiring users to rent the service directly.

A core technical differentiator for Sonic is its underlying architecture, running on state space models rather than conventional transformers. Cartesia frames traditional industry tradeoffs—such as speed versus naturalness and accuracy versus cost—as architectural challenges rather than inevitable limitations of the technology.

audio waveform digital sound mixing interface

Stock photo for illustration only, not from the actual event

In terms of practical output metrics, the platform focuses heavily on time-to-first-audio performance. Cartesia states a sub-90ms TTS latency alongside a 100ms transcript latency for its Ink-2 speech-to-text model. Both figures represent vendor-stated model latencies rather than measured end-to-end round trips.

<90msTTS Latency
$49.00Price per 1M Chars

Control features built into Sonic are tailored specifically for agent transcripts rather than standard narration tasks. Cartesia's launch demonstrations highlight English audio generation complete with natural pauses and filler words, alongside Hinglish code-switching capabilities between Hindi and English.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

The adoption of state space models (SSMs) instead of transformers marks a strategic shift for real-time streaming speech synthesis. SSMs excel at handling extended sequence lengths efficiently with lower memory overhead, enabling the ultra-low latency thresholds required for conversational AI agents and interactive voice applications.

Regarding market pricing, Artificial Analysis normalizes Sonic 3.6 at $49.00 per 1 million characters. This positions it at half the cost of ElevenLabs Eleven v3, which runs at $100.00, while sitting above Speechify Simba 3.2 at $10.00 for a 1,240 Elo rating.

Cartesia utilizes a credit-based sales model rather than direct character billing. The Scale tier, priced at $299 per month, includes approximately 10,667 TTS minutes and supports 15 concurrent requests. Dedicated line voice agents are billed separately at $0.06 per minute, with all figures verified as of August 18, 2026.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article