Google Releases Gemini 3.8 Live & Extended Thinking
Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking for production-grade voice agents, available now via Gemini Live API and Google AI Studio.

Stock photo for illustration only, not from the actual event
- Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for production voice agents.
- Standard model focuses on scale and cost efficiency, while Extended Thinking handles complex multi-step reasoning.
- Priced at $0.005/min for audio input and $0.018/min for audio output.
- Integrated with real-time partners including Agora, LiveKit, Salesforce, and Vercel.
Google has announced the official release of two new AI voice models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, engineered specifically to power production-grade voice agents via API. Both models are available starting today in the Gemini Live API and Google AI Studio. They operate as fully hosted models without open-weights availability, meaning self-hosting options are not provided.
The launch introduces two distinct models tailored for specific operational roles. Gemini 3.8 Live is optimized for scale and cost-efficiency, merging conversational intelligence with fluid dialogue and visual grounding capabilities. Meanwhile, Gemini 3.8 Live Extended Thinking targets high-complexity tasks, introducing enhanced intelligence and multi-step reasoning executed dynamically while speaking. Google frames both models as streamlined alternatives to traditional cascaded speech pipelines that chain separate ASR, LLM, and TTS components.
The Live API exposes several core capabilities across the new model lineup:
- Immediate availability within Gemini Live API and Google AI Studio ecosystems
- High-speed conversational fluidity for natural voice interactions
- Visual grounding integration built into the standard variant
- Dynamic multi-step reasoning while speaking in the Extended Thinking variant
- Seamless integration via real-time media streaming partners
Regarding pricing structure, both models are billed at $0.005 per minute for audio input and $0.018 per minute for audio output. Google clarifies that this estimate equates to $3 per 1 million input tokens and $12 per 1 million output tokens. Developers can leverage Live API integration partners that manage real-time media streaming infrastructure, including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents, alongside enterprise collaborations with Salesforce, Genspark, and Lumeris.

Stock photo for illustration only, not from the actual event
This release highlights Google's ongoing push to eliminate the latency bottlenecks inherent in traditional cascaded voice pipelines (ASR-to-LLM-to-TTS). By consolidating native multimodal reasoning into unified models, developers can build responsive voice agents capable of handling interruptions and complex logic without cumbersome middleware.
Developers can access technical documentation, developer posts, and functional example applications hosted directly on GitHub.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment