Google Brings Gemini Live Avatar to Enterprise
Google has made Gemini 3.8 Live with Live Avatar generally available, combining real-time dialogue, synchronized video personas, tools and 97-language support.
Google has made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise, turning its low-latency conversational model into a video-facing enterprise agent. The release combines bidirectional speech interaction with a generated on-screen persona that listens, responds and continues to appear engaged while backend tools run asynchronously.
What Google released
Live Avatar produces synchronized lip movement, facial expressions and natural turn-taking from audio and visual inputs. Google says the system can switch between 97 languages while preserving speech-to-video synchronization. Enterprises can choose from preset avatars, while custom avatar creation from a reference image is available through an allowlist.
The system is designed for customer-service desks, interactive kiosks, guided walkthroughs and other settings where a voice-only agent may feel too abstract. Its asynchronous tool execution is important: an avatar can continue the conversation while checking information or completing an operation in the background, rather than freezing during every external call.
Why the launch matters
This is a product-availability story, not merely a research demonstration. Google is placing a visually embodied agent inside Gemini Enterprise, where companies can connect it to existing workflows and deploy it across web, mobile and kiosk environments. That makes the competitive question less about whether AI can generate a talking face and more about whether organizations will accept synthetic presence as a frontline interface.
The release also exposes unresolved constraints. Custom likenesses remain gated, and the avatar’s usefulness depends on reliable tool execution, latency and organizational controls. Google says all generated audio and video carry SynthID watermarking, but watermarking does not by itself solve impersonation, disclosure or user-consent problems.
For the market, the consequential shift is that real-time multimodal agents are being packaged as an enterprise service with global language coverage, rather than left as an experimental demo.