Gradium AI Releases New TTS Model: 81.0% Hard-Case Pass Rate
Gradium AI rolled out its new default TTS model on API and Studio on August 31, 2026, achieving an 81.0% hard-case pass rate at 216 ms first audio.

Stock photo for illustration only, not from the actual event
- Gradium AI launched its new default TTS model across API and Studio on August 31, 2026, with zero migration required.
- Achieved an 81.0% hard-case pass rate across a 500-sentence evaluation set over five languages and ten criteria.
- Delivered a P50 time-to-first-audio of 216 ms with the tightest latency variance among tested models.
- Open-sourced the 500-sentence evaluation benchmark on Hugging Face under the CC BY 4.0 license.
Gradium AI has officially rolled out a major update, deploying its brand-new text-to-speech (TTS) model as the default option across its API and Studio platforms starting August 31, 2026. The transition is designed to happen seamlessly with no migration needed, allowing all existing voices and custom voice clones to continue functioning without any manual adjustments.
To establish rigorous benchmarks, Gradium built a 500-sentence evaluation set and open-sourced it on Hugging Face under a CC BY 4.0 license. The dataset spans five languages—English, German, French, Spanish, and Portuguese—and evaluates models across ten criteria. These include seven atomic criteria covering spelling, acronyms, alphanumeric tokens, dates, regular numbers, large/floating numbers, and email, alongside three composite criteria tracking orders, IT tickets, and claims.
Scoring was conducted through strict human evaluation, requiring independent native-speaker raters to verify that every single phonetic element was pronounced accurately. A single dropped digit resulted in an immediate failure for the sentence. Pooled across all ten criteria and averaged equally across the five languages, Gradium TTS scored 81.0%, outperforming Cartesia Sonic 3.6 at 75.1%, ElevenLabs v3 Conversational at 65.4%, Fish Audio S2.1 Pro at 49.5%, and Inworld TTS 1.5 Max at 46.5% using default settings in August 2026.

Stock photo for illustration only, not from the actual event
On Coval’s TTS benchmark, Gradium reported a P50 time-to-first-audio of 216 milliseconds, representing a 170 ms improvement over the model it replaces. A more critical metric is the interquartile range (p75-p25 spread) of 30 ms across 480 runs, marking the tightest variance among the five models tested. By comparison, Cartesia Sonic 3.6 recorded a 454 ms median with a 165 ms spread.
Focusing on latency spread rather than just median speed is crucial for conversational AI applications, as end-users are often impacted by tail latency spikes rather than average figures. Gradium's ability to maintain a minimal 30 ms variance helps eliminate unpredictable audio lag and jitter during live voice agent interactions.
While Inworld TTS 2 achieved a faster median time of 166 ms, and models like Fish Audio S2.1 Pro (291 ms) and ElevenLabs v3 Conversational (329 ms) trailed behind, Gradium's competitive edge lies in achieving the lowest hard-case error rate while maintaining sub-250 ms latency with minimal variance.
Existing users can continue utilizing the system without interruption, while new engineering teams can adopt the model by installing the Python SDK and pointing to the WebSocket TTS endpoint using existing voice IDs. Gradium is also offering 1 million credits on its Discord for complete hard-case failure reports.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment