SpaceXAI Launches Grok Voice Transcribe 2.0 API with 2x Accuracy
SpaceXAI releases Grok Voice Transcribe 2.0 speech-to-text API offering double the accuracy of v1.0 at $0.10 per hour, adopted by Atlassian Loom.

Stock photo for illustration only, not from the actual event
- SpaceXAI debuts Grok Voice Transcribe 2.0 hosted API under model ID grok-voice-transcribe-2.0.
- Delivers twice the accuracy of version 1.0 with a sharp drop in word error rate (WER) for short phrases.
- Pricing remains fixed at $0.10 per hour for batch processing and $0.20 per hour for streaming.
- Atlassian Loom integrates the new API for video transcription alongside Cursor.
SpaceXAI has officially unveiled its latest speech-to-text model, Grok Voice Transcribe 2.0, now live as a hosted API under the model ID grok-voice-transcribe-2.0. The model builds upon the audio foundation model powering Grok Voice, a system that already manages tens of thousands of customer support calls daily, transcribes millions of hours of video, and operates the Grok assistant inside Tesla vehicles.
The training pipeline utilized live, noisy, and multilingual audio gathered across diverse environments, followed by post-training refinements. According to SpaceXAI, the model secured a first-place ranking among 32 streaming models on the public Artificial Analysis leaderboard (AA-WER Streaming), which evaluates roughly 8 hours of audio data.

Stock photo for illustration only, not from the actual event
Regarding multilingual capabilities, version 2.0 automatically detects languages and handles mid-recording language switches within a single pass. SpaceXAI highlights multilingual accuracy as its biggest upgrade over version 1.0. For short phrases such as in-car commands, where contextual clues are minimal, the word error rate (WER) drops significantly from 20.6% down to 6.8%, translating to roughly 67% fewer errors. Furthermore, documentation lists support for written-form formatting across 25 languages.
The introduction of SpaceXAI's upgraded transcription API highlights the industry push toward production-ready speech recognition that tackles edge cases like brief voice commands and rapid code-switching. By optimizing performance on noisy, real-world audio streams, the model addresses common failure points in multi-speaker and multilingual environments without requiring cumbersome manual configurations.
Features included natively in the API comprise a batch endpoint accepting files up to 500 MB across 12 audio formats, alongside streaming support for Opus at approximately 4 KB/s compared to 48 KB/s for raw PCM at 24 kHz.
"closing the loop from context to code."
Sanchan Saxena, SVP of Teamwork Collection at Atlassian
Pricing structures remain unchanged from version 1.0. Batch transcription costs $0.10 per hour of audio, while streaming is priced at $0.20 per hour, equating to approximately $1.67 and $3.33 per 1,000 minutes respectively. Diarization, timestamps, and key terms are bundled at no extra cost.
In practical applications, Atlassian Loom has integrated Grok Voice Transcribe 2.0 to handle video transcriptions after finding it superior to prior solutions. The workflow involves recording an action plan in Loom and piping the resulting transcript directly into Cursor for automated code updates. Sanchan Saxena, SVP of Teamwork Collection at Atlassian, described the integration as closing the loop from context to code.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment