Live Avatar builds on last week's Gemini 3.8 Live launch: it processes visual and audio input simultaneously and replies in near real time with speech plus video, including precise lip-sync, natural expressions and fluid turn-taking.
Tool calls run asynchronously: the avatar can fetch data in the background while the conversation continues, which Google demonstrates with a hotel guest check-in flow.
Google says the avatar can switch among 97 languages mid-conversation, adapting lip-sync and expressions without degrading video fidelity or causing visual drift.
Beyond a preset library, developers can generate a custom animated avatar from one high-quality reference image, but custom avatar creation currently requires enterprise allowlisting. All audio and video output carries a SynthID watermark.
Terms in this piece · Glossary
streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
Why it matters
Gemini Enterprise now pairs near real-time video generation with speech for lip-synced, expressive virtual personas, letting teams build interactive avatar-based support or walkthrough experiences without stitching together separate video and audio models.