Gradium Launches Voice Design for Synthetic Voices
Paris-based voice AI startup Gradium has launched Voice Design, allowing users to generate brand new synthetic voices from text prompts within seconds.

Stock photo for illustration only, not from the actual event
- Gradium introduces Voice Design to generate new voices from text descriptions of 1 to 500 characters.
- The process supports five languages including English, French, Spanish, Portuguese, and German.
- The system generates 1 to 5 candidate variations per request within 3 to 5 seconds.
- Blind pairwise listening tests report a 72.6% win rate for Gradium against six competing systems.
Paris-based voice AI company Gradium, spun out of the Kyutai research lab, has released Voice Design, a tool enabling users to create entirely new synthetic voices within seconds using only written descriptions. The system operates without requiring any reference audio, speaker models, or rights clearance.
Voice Design is currently live within the Gradium API and Studio, offered for free across all tiers including the free plan. Saved voices run on the same streaming Text-to-Speech endpoint as catalog voices, maintaining identical latency and output formats. Users can input descriptions ranging from 1 to 500 characters in length.

Stock photo for illustration only, not from the actual event
The text description serves as the sole input for the model. Gradium's documentation outlines specific attributes the system responds to, such as gender, age band, accent or origin, pitch, pace, energy, timbre, resonance, register, manner, and the intended job of the voice. Gradium advises users to conclude descriptions with the intended use to effectively guide delivery and register.
A single request yields 1 to 5 candidate variations, typically prepared within 3 to 5 seconds. These represent variations of a single character, meaning distinct characters require unique descriptions rather than additional samples. The workflow involves four distinct API calls, spanning from candidate generation and embedding retrieval to promoting the chosen voice.
Gradium's prompt-based voice generation approach significantly streamlines the workflow for developers and content creators by removing the overhead of recording and licensing human voice talent. While this offers unprecedented speed in character creation, ensuring long-form consistency and emotional depth across varied prompts remains a key technical frontier for synthetic voice generators.
"Gradium ran a blind pairwise listening test on accent prompts across six voice design systems reachable through public APIs and five languages."
MarkTechPost
To evaluate performance, Gradium conducted a blind pairwise listening test comparing six voice design systems across five languages over 7,627 comparisons. Gradium achieved a 72.6% win rate, outperforming ElevenLabs at 59.0%, Inworld at 44.8%, Fish Audio at 36.7%, and MiniMax at 31.7%. Additionally, Gemini 3.1 Pro rated single unlabelled clips and placed Gradium at the top with a score of 4.06.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment