Google Gives Gemini a Face: Gemini 3.8 Live Avatar With Real-Time Lip-Sync in 97 Languages

Google's Gemini 3.8 Live Avatar adds a real-time talking face with lip-sync across 97 languages to Gemini Enterprise — what it means for creators and brands.

Google Gives Gemini a Face: Gemini 3.8 Live Avatar With Real-Time Lip-Sync in 97 Languages
Table of contents

Google Just Gave Gemini a Face — and It Speaks 97 Languages in Real Time

On September 24, 2026, Google introduced Gemini 3.8 Live with Live Avatar, and made it available immediately for business customers through Gemini Enterprise, following up the next day with a general availability post on the Google Cloud blog. The idea in one line: after Gemini 3.8 Live — which we covered last week — learned to listen and speak in Arabic and 96 other languages, this release adds the eyes: a visual persona that talks with you in near real time, with live lip-syncing and mid-conversation language switching that does not break the video.

Official image from Google's Gemini 3.8 Live Avatar announcement
Official artwork from Google's announcement (source: Google blog)

What Exactly Is Live Avatar?

According to the official announcement, the feature couples "live dialogue capabilities with low-latency streaming video" to produce an avatar that listens, sees, and speaks in near real time. Three qualities define the experience, in Google's own words: precise lip-syncing, natural expressions, and fluid turn-taking. The technology was first previewed at Google Cloud Next 2026 and became generally available for enterprise production this week.

It runs on the web, on mobile, and on interactive kiosks — and it understands what you show it: the model processes live camera feeds and screen shares alongside audio, so it can look at what you are filming or presenting and talk about it directly.

Four Capabilities That Matter to Creators and Brands

  • 97 languages with automatic detection: the model understands and speaks 97 languages, and its lip-sync and expressions adapt dynamically when you switch languages mid-conversation "without degrading video fidelity or introducing visual drift" — Arabic is one of the core Live family languages.
  • Background tool calling without going silent: it can trigger API calls and fetch data while the dialogue continues uninterrupted — Google demonstrated a hotel guest check-in handled exactly this way.
  • Natural interruption recovery: the native speech-to-speech foundation recovers conversation context even if the user interrupts mid-transaction.
  • Custom avatars from a single photo: from one high-quality reference image plus an audio sample, developers can generate a fully animated, responsive avatar that preserves the reference likeness and brand styling. This is currently gated behind an enterprise allowlist, while the library of diverse preset avatars works for everyone.
Official showcase of Gemini 3.8 Live avatars
Official avatars from Google's announcement, each with a distinct look and voice (source: Google blog)

What Does This Mean for You as a Creator or Brand Owner?

Until today, the "virtual presenter" has been a recorded-video product: write a script, generate the video, publish. What changed this week is the arrival of the interactive presenter: a face for your brand that greets visitors on your site, in your app, or at your event kiosk, answers in your language, sees what is shown to it, and calls your systems as needed — commercially available from Google at production grade for the first time.

Three scenarios sit close to our audience: visual customer service for an online store (the avatar sees the product you hold up to the camera and explains it), interactive guides inside education or real-estate apps, and event kiosks that welcome visitors without extra staff. And when you need recorded content instead of live interaction, note that Google's own launch partners for the companion Gemini 3.8 Flash TTS release include HeyGen itself — the infrastructure under your avatar tools is consolidating, which usually means prices follow.

For publishers of regular content — videos, articles — the benefit today is indirect: the cost of a "talking face" is falling and its quality is maturing, while the quality of the script underneath remains the decisive factor in whether anyone watches. That is exactly where a writing-first toolchain such as ARWriter fits: get the script right before you put a synthetic face on top of it.

Transparency Is Built In: SynthID on Every Frame

Every audio and video stream the avatar produces carries an imperceptible SynthID watermark woven directly into the generated content, keeping AI output detectable. Identity rules are explicit alongside it: a curated library of preset avatars, and custom avatar creation wrapped in verification and an enterprise allowlist. For anyone worried about voice-and-face impersonation, the design is reassuring on both ends: the face speaking for your brand can be proven to be generated, and a face you do not own cannot be generated at all.

Quick Comparison: Where It Stands Today

DimensionGemini 3.8 Live AvatarRecorded avatar tools (e.g., HeyGen)
ModeLive two-way conversationRecorded video from a script
Live visionCamera feed + screen share, in real timeNot available (fixed script)
Languages97, auto-detected, instant switchingVaries by tool
AvailabilityGemini Enterprise (US/EU endpoints)Instant public subscriptions
PricingEnterprise contracts (no public rates)Published monthly plans
Live capture of the Google Cloud general availability post for Live Avatar
Live capture of the Google Cloud general availability post, September 25, 2026

Where to Start If Your Organization Is Ready

Official availability runs through Gemini Enterprise, and Google's technical documentation points to two paths: trying the model and its features from the Agent Platform Studio console (the Multimodal Live surface), and building programmatically against the Live API documentation on the same platform. Before a first conversation with sales, prepare answers to three questions that will shape the engagement: where the avatar will live (web, app, or kiosk), which languages your audience actually needs — and whether mid-session switching matters — and whether you require a custom avatar carrying your brand identity, because that alone activates the allowlist and verification track.

One practical tip from covering voice-interface tools all month: start with one tightly scoped pilot — a single FAQ page or a single reception scenario — and measure real-world latency and how the avatar behaves when interrupted repeatedly in your audience's actual dialect before scaling out. The platform ships with US and EU endpoints, provisioned throughput, enterprise compliance, and data governance, but the first pilot is still the only honest test of whether your audience will embrace a synthetic face.

Limitations You Should Know

  • Enterprise first: available only through Gemini Enterprise, with US and EU endpoints — no individual plan yet.
  • No public pricing: contracts only; no per-minute or per-session rates published.
  • Custom avatars are gated: allowlist plus verification, so do not expect to clone any face you like.
  • Extended Thinking is not here yet: that variant of Live remains in private preview.
  • Permanent watermarking: great for transparency, and it structurally closes the door on "undetected deepfake" misuse — which is precisely the point.

The Bottom Line

In the space of one week, Google completed the "voice and face" loop: a live conversation model that speaks Arabic, then a live face with lip-sync across 97 languages, and — as we explained in our guide to Gemini 3.8 Flash TTS — a full voice studio for recorded content. Anyone planning an Arabic-first channel or a talking storefront has more leverage than ever, and the lasting competitive edge will still belong to the quality of the content itself. Draft your script first with ARWriter's writing tools, then pick the face that delivers it.

Official Sources

Frequently Asked Questions

Does Gemini 3.8 Live Avatar support Arabic?

The model is built on Gemini 3.8 Live, which understands and speaks 97 languages with automatic language detection, and Arabic is among the supported Live family languages; lip-sync adapts automatically when switching languages within a single conversation.

Can I use it as an individual creator?

Current general availability is through Gemini Enterprise for organizations, with US and EU endpoints, and custom avatars require an enterprise allowlist; no individual plan has been announced.

How is it different from recorded AI presenter videos?

Those are pre-generated videos from a fixed script. Live Avatar is a live two-way conversation that sees the user's camera and screen and can call external tools mid-dialogue without stopping.

Can the generated video be detected as AI-made?

Yes. Every audio and video output carries Google's imperceptible SynthID watermark woven into the content itself, keeping it detectable and attributable.

How does it relate to Gemini 3.8 Flash TTS announced a day earlier?

Both belong to the Gemini Audio family: Flash TTS turns scripts into performed voice for recorded content such as podcasts and audiobooks, while Live Avatar is a live visual interface for interactive enterprise conversations.