Gemini 3.8 Live and Extended Thinking: Google’s Real-Time Voice AI Explained

Google’s new voice models think while they talk: where to try them, what the independent benchmarks say, and three experiments creators can run today.

Gemini 3.8 Live and Extended Thinking: Google’s Real-Time Voice AI Explained
Table of contents

Google has launched two new real-time voice models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, calling them its most advanced live dialogue models yet. The announcement, published on the official Google blog on September 15, 2026, brings parallel reasoning to voice conversations, automatic switching between 97 languages mid-sentence, and background task execution that keeps the chat flowing while tools run — available immediately across the Gemini app, Search Live, Google Workspace, and the Gemini API.

Primary source: Google’s official Gemini 3.8 Live announcement, September 15, 2026.

One number frames the launch: Extended Thinking took the number-one spot on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6, ahead of every competing model on the independent leaderboard. For anyone who talks to their tools — podcasters, marketers, developers brainstorming on the move — this launch is about removing the stop-and-wait rhythm that has defined voice assistants so far.

Official Google announcement graphic for Gemini 3.8 Live
Official announcement image — Source: Google blog, September 15, 2026

Two models, two jobs

  • Gemini 3.8 Live is built for scale and cost efficiency: fluid conversation plus real-time visual grounding, meaning it processes what your camera or screen shows while you keep talking.
  • Gemini 3.8 Live Extended Thinking is the heavyweight, aimed at high-complexity, multi-step reasoning. Its signature capability is thinking and speaking simultaneously — it acknowledges requests with natural early verbal cues like "Let me check that…", then narrates live progress while tools and API calls finish in the background.

In practice, the first is a fast conversationalist; the second is a voice agent that can coordinate bookings, function calls, and background research without going silent on you.

The official numbers, straight from the primary source

  • #1 on Artificial Analysis' Speech-to-Speech Quality Index (82.6) for Extended Thinking.
  • Leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking benchmark.
  • 97.7% on Big Bench Audio for audio reasoning.
  • Gemini 3.8 Live took second place on Speech Agent Arena based on user preferences.
  • On ServiceNow's EVA-Bench for voice agents, Google says the models push the accuracy-versus-conversational-quality frontier forward.

These figures come from independent evaluators (Artificial Analysis, Speech Agent Arena), though the selection of benchmarks is Google's own — healthy skepticism remains a feature, not a bug.

Where you can use the new models today

  • Everyone: Gemini 3.8 Live powers Search Live starting on launch day.
  • Extended Thinking for everyone: in the Gemini app's Live mode, plus Gmail and Keep for all Google AI subscribers, and Docs for Google AI Pro and Ultra subscribers in Workspace.
  • Developers: both models are live in the Gemini API and Google AI Studio from day one.
  • Enterprises: private preview in Gemini Enterprise, with the customer experience platform coming soon.

The developer ecosystem moved fast too: platforms like Agora, Fishjam, LangChain, LiveKit, Pipecat, and Vercel already support the Live API, and Salesforce, Genspark, and Lumeris are early adopters. Expect a wave of voice-first apps built on these models in the coming weeks.

Screenshot of the official Gemini 3.8 Live announcement page on the Google blog
The official announcement page on blog.google, September 15, 2026

Why this matters if you make content for a living

Voice-first drafting without the dead air. Ask Extended Thinking to outline a week of content and audit your recent headlines, and it acknowledges instantly, works in the background, and walks you through results as they land. Google's live demos included building complete business plans and custom marketing toolkits through speech alone.

Eyes while it talks. Near-real-time visual input means you can point your camera at a design draft, a thumbnail, or a storyboard and discuss it out loud — a natural fit for design reviews and presentation practice.

97 languages, switched automatically mid-conversation. The model detects and transitions between supported languages on the fly — useful for bilingual creators, localization reviews, and interviews that drift between languages. Multilingual creators who need polished written output afterward can pass the transcript to a dedicated writing platform such as ARWriter's toolset for editing and publishing.

SynthID is stamped on everything. All audio generated by Google's AI products carries the imperceptible SynthID watermark, making AI audio detectable. If you produce voiceovers or podcast segments with these models, know that the output is watermarked by design — a transparency measure the industry is standardizing on.

Quick comparison

  • Versus Gemini 3.8 Flash (the text model): Flash handles fast written work; Live handles spoken, real-time interaction. They complement rather than replace each other — we covered the Flash line separately when Gemini 3.8 Flash launched earlier in September.
  • Versus ChatGPT's voice mode: Google's claim to the lead rests on the independent Speech-to-Speech index and on background tool execution while speaking; actual feature availability varies by plan and platform, so test both on your own workload.
  • Versus traditional phone assistants: the gap is no longer command recognition — it is continuity: background tasks, live vision, and multilingual switching inside a single session.

Three experiments to start with today

You do not need a complex plan to measure what the new capabilities do for your work — here are three experiments, each under fifteen minutes:

Experiment 1 — real-time voice note-taking. Open Live mode in the Google app and talk through your next piece: why you are writing it, who it is for, what is new. Ask for the key points as a list at the end. With automatic language switching you can drop English terms inside your sentences and the context holds. The output: a structural draft ready for editing.

Experiment 2 — visual design review. Point your camera at a thumbnail you are considering and ask: which element draws the eye first? Is the text readable from a distance? Real-time visual grounding means you and the model are discussing the same thing — free training for your design eye.

Experiment 3 — a background task while you keep talking. Start a conversation and give Extended Thinking a slow job — collecting headline ideas for ten posts, grouped by angle — then continue discussing something else. Notice how it acknowledges the request instantly and returns with the result narrated as progress. That simultaneity is the fundamental difference from earlier assistants that went silent while working.

Honest limitations

  • The strongest model is partially gated: Extended Thinking in Docs requires Google AI Pro or Ultra, and enterprise availability is still a private preview.
  • Benchmark leadership is real but measured on Google's chosen metrics; day-to-day differences between top models are often smaller than leaderboard gaps suggest.
  • SynthID watermarking means generated audio is identifiable — great for transparency, restrictive for anyone hoping to pass synthetic audio as human-recorded.
  • Long-form quality and accent handling across the 97 languages were not detailed in the announcement; validate on your own language and use case before committing.

Frequently asked questions

What is Gemini 3.8 Live?

It is Google's new real-time voice model announced on September 15, 2026, alongside a stronger variant called Gemini 3.8 Live Extended Thinking. It supports near-real-time visual input, background task execution, and automatic switching between 97 languages during a conversation.

What is the difference between Gemini 3.8 Live and Extended Thinking?

Gemini 3.8 Live is the fast, cost-efficient conversational model, while Extended Thinking targets complex multi-step tasks and reasons while speaking, narrating progress as it completes background work.

Is Gemini 3.8 Live free to use?

Gemini 3.8 Live is available to everyone in Search Live. Extended Thinking is available in the Gemini app and in Gmail and Keep for Google AI subscribers, while Docs access requires Google AI Pro or Ultra in Workspace.

Which languages does Gemini 3.8 Live support?

The model automatically detects and transitions between 97 supported languages mid-conversation, including widely used languages such as English, Arabic, Spanish, French, and Hindi, per the official Live API language list.

Can developers build with Gemini 3.8 Live?

Yes. Both models are available in the Gemini API and Google AI Studio from launch day, with partner platforms like LiveKit, LangChain, Agora, Pipecat, and Vercel supporting real-time voice interfaces built on the Live API.

Bottom line

Gemini 3.8 Live shifts voice AI from "answers quickly" to "works while you talk": live vision, 97 languages, background execution, and independent benchmark leadership. The fastest win for creators is folding it into an existing pipeline — think out loud, capture the output, then turn spoken ideas into polished written content with a dedicated multilingual writing tool like ARWriter, from spoken idea to publishable post.