On September 24, 2026, Google introduced Gemini 3.8 Live with Live Avatar on its official blog: a near real-time video layer on top of the live dialogue model launched the week before, producing a lip-synced, expressive persona that listens, sees and speaks. The post came from Google DeepMind research scientist Shuo-yiin Chang and Gemini software engineer CJ Zheng, and the feature is generally available in Gemini Enterprise starting today (official announcement).
The face is generated by the model itself
The technically interesting part is "native". Live Avatar is not a voice model feeding a separate animation engine: per Google's developer documentation (verified by WindowsForum), 3.8 Live outputs a 24 FPS MP4 video stream synchronized with synthesized speech through the Live API — developers request video as an output type (response_modalities=["VIDEO"]) the same way they request audio, and the face is produced by the model itself. Visual and audio inputs are processed simultaneously, keeping mouth movement, expression and content aligned in near real time.
97 languages and background tool calls
The feature ships native multilingual speech-to-speech synchronization: mid-conversation language switching adapts lip-sync and expressions across 97 languages without degrading video fidelity or introducing visual drift. Another engineering highlight is asynchronous tool execution — the avatar can trigger tool calls and fetch data in the background while the dialogue continues; Google's demo scenario is a hotel check-in where the agent confirms details with the guest while completing the operation via APIs. Google's Cloud blog adds that the native speech-to-speech foundation recovers from interruptions naturally, without dropping conversation context or backend transactions (GA post).
Enterprise first, watermarks and allowlists included
Live Avatar currently targets Gemini Enterprise customers (US and EU endpoints); consumer Gemini apps don't have it yet. Enterprises get a library of preset avatars and can generate custom ones from a high-quality reference image that preserves likeness — through enterprise allowlisting. All audio and video output carries the invisible SynthID watermark to keep AI-generated content detectable. The sibling 3.8 Live Extended Thinking remains in private preview.
So what
The Register noted the awkward timing: Google's own researchers have spent years warning that humanlike AI leads people to trust it more than they should — and the enterprise product has already shipped (report). Technically, the route compresses the two-stage "voice model plus animation engine" pipeline into end-to-end multimodal output — simpler to deploy, shorter latency path, and language consistency guaranteed by the model itself. For enterprises this is a strong candidate default for customer service and kiosk scenarios; for everyone else, the detail worth watching is the invisible watermark: as on-screen faces look more human, a verifiable "this is AI" mark may be the last thing we can hold onto.