Google says Live Avatar couples Gemini 3.8 Live's dialogue model with near real-time video generation, producing lip-synced, expressive video while it processes audio and visual input at the same time.
Asynchronous tool useA model's ability to call external functions — run code, search the web, edit files — instead of only generating text.Full definition → lets the avatar fetch data in the background while it keeps talking; Google's demo shows a hotel guest check-in handled without pausing the conversation.
Google claims lip-sync and expressions adapt across 97 languages, including switches mid-conversation, without degrading video fidelity or introducing visual drift.
Beyond preset avatars, developers can animate a custom avatar from one high-quality reference image, but only via enterprise allowlisting; all audio and video output carries a SynthID watermark.
Terms in this piece · Glossary
tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
Why it matters
Pairing low-latency video generation with live dialogue changes what's possible for enterprise conversational agents, moving past voice-only interfaces toward visually embodied assistants.