Google launches Gemini 3.8 Live voice models
Google’s Gemini 3.8 Live models target production voice agents with faster responses, visual understanding, 97-language switching, and lower operating costs.
Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of native audio models aimed at production-grade voice agents rather than simple conversational demos. The models are rolling out to consumers through Search Live, to developers in public preview through the Gemini API and AI Studio, and to enterprises through private preview channels.
Built for continuous interaction
Google says both models combine conversational reasoning with near-real-time audio and visual understanding. They are designed to listen, decide when to respond, maintain context through interruptions, and execute tasks during a conversation. The Extended Thinking version adds a slower reasoning mode for cases where the agent must plan before speaking or acting.
The launch also emphasizes multilingual interaction. Google says the models can switch between 97 languages, a capability that matters for customer-service and travel scenarios where users may move between languages in the same session. The release is accompanied by a model card covering intended uses, evaluation results, safety mitigations, and distribution terms.
A cost and latency contest
Independent measurements cited in the launch cycle place Gemini 3.8 Live at roughly $0.84 per hour of input audio, compared with about $1.75 for Gemini 3.1 Flash Live. Artificial Analysis also reported a 1.18-second average time to first audio on its Big Bench Audio benchmark, with the Extended Thinking configuration at 1.35 seconds. Those figures, if sustained in production, make the model more relevant to high-volume call handling than a model that is merely strong in isolated conversations.
The release matters because voice agents are moving from a feature layer into an operating layer for support, sales, tutoring, and task execution. Lower audio cost reduces the penalty for keeping conversations open, while lower latency makes interruptions and turn-taking feel less mechanical. The open question is whether benchmark speed and language coverage survive noisy environments, long calls, tool failures, and business-specific safety constraints. Google is now competing on the full economics of spoken interaction, not only on model intelligence.