Google launches Gemini 3.8 Live voice models with background reasoning
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, introducing two native-audio models for conversations that can continue while the system reasons and uses tools. The announcement is a product release rather than a research preview: Google says both models are rolling out through the Gemini API and Google AI Studio, while enterprise and consumer availability differs by product and subscription.
The standard Gemini 3.8 Live is designed for scale and cost efficiency. Gemini 3.8 Live Extended Thinking is aimed at more complex tasks that require multi-step reasoning. Both models can process audio, images, video and text, and return audio and text. Google says Live can understand visual context in near real time, switch automatically between 97 languages during a conversation, and make tool or API calls in the background without forcing the user to wait in silence.
The extended-thinking model adds a different interaction pattern. It can reason and speak at the same time, using short progress cues while an asynchronous task continues. Google gives examples involving live troubleshooting, business planning, coding and multi-step bookings. The company says the models are available to developers through the Gemini API and AI Studio; Live is also rolling out in Search Live, and the extended model is appearing in Gemini Live and selected Gmail, Docs and Keep experiences.
Google reports that Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’s Speech to Speech Quality Index, 68.6 percent on the τ-Voice agentic-performance benchmark and 97.7 percent on Big Bench Audio. Google also says standard Live ranked second in the Speech Agent Arena. These are measured results from named evaluation systems, but the launch article’s benchmark selection and deployment setup still matter; they are not a guarantee for every voice workflow.
The practical change is that a voice agent can acknowledge a request, keep the conversation moving and finish work in parallel. That could make voice more useful for customer service, accessibility, field work and hands-free productivity. It also raises the importance of permissions, interruption handling and clear disclosure when a model acts in the background. Google says generated audio carries its SynthID watermark. The model card still notes ordinary foundation-model limitations such as hallucinations, slowness and timeouts, so the release expands what voice agents can attempt without removing the need for human review.