Skip to main content

Gradium Launches Voice Design for Synthetic Voices

Paris-based voice AI startup Gradium has launched Voice Design, allowing users to generate brand new synthetic voices from text prompts within seconds.

AI-written
Inewgen
09 Sep 2026Source: MarkTechPost3 min read (0 views)
Share
Gradium Launches Voice Design for Synthetic Voices

Stock photo for illustration only, not from the actual event

Font size
  • Gradium introduces Voice Design to generate new voices from text descriptions of 1 to 500 characters.
  • The process supports five languages including English, French, Spanish, Portuguese, and German.
  • The system generates 1 to 5 candidate variations per request within 3 to 5 seconds.
  • Blind pairwise listening tests report a 72.6% win rate for Gradium against six competing systems.

Paris-based voice AI company Gradium, spun out of the Kyutai research lab, has released Voice Design, a tool enabling users to create entirely new synthetic voices within seconds using only written descriptions. The system operates without requiring any reference audio, speaker models, or rights clearance.

Voice Design is currently live within the Gradium API and Studio, offered for free across all tiers including the free plan. Saved voices run on the same streaming Text-to-Speech endpoint as catalog voices, maintaining identical latency and output formats. Users can input descriptions ranging from 1 to 500 characters in length.

chromebook notebook computer office desk workspace

Stock photo for illustration only, not from the actual event

The text description serves as the sole input for the model. Gradium's documentation outlines specific attributes the system responds to, such as gender, age band, accent or origin, pitch, pace, energy, timbre, resonance, register, manner, and the intended job of the voice. Gradium advises users to conclude descriptions with the intended use to effectively guide delivery and register.

72.6%Gradium Win Rate
13.6%Lead Over ElevenLabs

A single request yields 1 to 5 candidate variations, typically prepared within 3 to 5 seconds. These represent variations of a single character, meaning distinct characters require unique descriptions rather than additional samples. The workflow involves four distinct API calls, spanning from candidate generation and embedding retrieval to promoting the chosen voice.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Gradium's prompt-based voice generation approach significantly streamlines the workflow for developers and content creators by removing the overhead of recording and licensing human voice talent. While this offers unprecedented speed in character creation, ensuring long-form consistency and emotional depth across varied prompts remains a key technical frontier for synthetic voice generators.

"Gradium ran a blind pairwise listening test on accent prompts across six voice design systems reachable through public APIs and five languages."

MarkTechPost

To evaluate performance, Gradium conducted a blind pairwise listening test comparing six voice design systems across five languages over 7,627 comparisons. Gradium achieved a 72.6% win rate, outperforming ElevenLabs at 59.0%, Inworld at 44.8%, Fish Audio at 36.7%, and MiniMax at 31.7%. Additionally, Gemini 3.1 Pro rated single unlabelled clips and placed Gradium at the top with a score of 4.06.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article