Skip to main content

Kyutai Releases Voice of Reason Speech Math Model

Kyutai has launched Voice of Reason, a speech-native model solving spoken math with reinforcement learning, achieving up to 77.1% on GSM8K.

AI-written
Inewgen
23 Sep 2026Source: MarkTechPost2 min read (0 views)
Share
Kyutai Releases Voice of Reason Speech Math Model

Stock photo for illustration only, not from the actual event

Font size
  • Kyutai releases a speech-native model solving math via RL
  • Interleaves 13 text tokens and 26 audio tokens in output
  • Achieves benchmark scores up to 77.1% on GSM8K
  • Trained on 16 H100 GPUs with 1,500 RL updates

AI research organization Kyutai has announced the release of Voice of Reason, described as the first application of reinforcement learning to mathematical reasoning within a speech-native model, designed to bypass traditional pipeline limitations.

While cascaded pipelines involving speech-to-text, text LLMs, and text-to-speech still lead in general reasoning, each intermediate stage introduces latency and drops paralinguistic cues such as tone. Speech-native models must generate audio at regular intervals to maintain interactivity, placing constraints on their internal reasoning tokens.

77.1%Top GSM8K Score
1,500RL Training Updates

The base GLM-4-Voice model initially scored 27.3% on the GSM8K benchmark, which was improved to 58.7% using the STITCH method. The new architecture interleaves its outputs systematically, generating 13 text tokens followed by 26 audio tokens in a repeating sequence.

ai research laboratory computer hardware technology

Stock photo for illustration only, not from the actual event

Training utilized a group-relative REINFORCE objective related to GRPO while omitting PPO clipping and KL regularization. The process was executed across 16 H100 GPUs over 1,500 reinforcement learning updates.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Kyutai's shift toward a native speech reasoning model highlights the industry's focus on reducing latency for real-time voice applications. By embedding reinforcement learning directly into a speech-to-speech workflow, the architecture paves the way for voice assistants capable of complex logical deduction without relying on intermediate text conversions.

Evaluation scores utilizing top-k 50 decoding averaged across three seeds reached 70.3% and 77.1% when top-k was removed. The released model checkpoints are now available for self-hosting on H100 infrastructure under the GLM-4-Voice license.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article