Google Launches Gemini 3.5 Transcribe: AI Speech-to-Text That Cleans Your 'Ums' and Labels Speakers

Google's new Gemini 3.5 Transcribe turns raw audio into polished, formatted text — automatic filler removal, three-speaker attribution, and word-level timestamps.

Google Launches Gemini 3.5 Transcribe: AI Speech-to-Text That Cleans Your 'Ums' and Labels Speakers
Table of contents

On August 26, 2026, Google introduced Gemini 3.5 Transcribe, a new speech-to-text model the company describes as its most precise yet. The detail that matters to anyone who works with audio is simple: instead of producing a raw, literal transcript, the model converts rough audio directly into accurate, polished, formatted text — cutting filler words like "um" and "uh," picking up specialized jargon, attributing speakers, and stamping every single word with a timestamp.

The announcement came from Google's Gemini Audio team on the official company blog. Developers can already try the model in public preview through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, while consumer availability has started gradually: the Gemini app on macOS and the Rambler dictation feature on Android in select countries and languages, with Chrome support promised "soon."

Official hero image for the Gemini 3.5 Transcribe announcement on Google's blog
Official image from Google's blog (blog.google), August 26, 2026

What Gemini 3.5 Transcribe actually does

The model is built around three connected jobs that every audio-first creator runs into: transcribe, clean, and understand. Rather than handing you a messy wall of text that needs an hour of cleanup, it returns publication-ready output. The headline capabilities include:

  • Automatic disfluency removal: hesitation sounds and repeated words disappear from the final text without manual editing.
  • Speaker attribution: in pre-recorded audio it can attribute speech for up to three separate speakers.
  • Word-level timestamps: essential for subtitled videos, translated captions, and clip hunting inside long recordings.
  • Custom vocabulary: feed it your brand names and niche terms, and the model respects your spelling instead of guessing.
  • Live language switching: it handles speakers who mix two languages mid-sentence — a daily reality for bilingual creators — with seamless streaming transcription.
  • Smart formatting: output arrives structured with paragraphs and headings rather than one unbroken block.
Google's official side-by-side demo: conventional transcription versus Gemini 3.5 Transcribe (source: blog.google)

The performance numbers Google published

Google positions the new model as a major step up from its previous transcription model, Chirp 3, on three fronts: accuracy, latency, and entirely new capabilities. According to measurements by Artificial Analysis, time to final transcription improves by 70%. On the multilingual FLEURS benchmark, the model scores a 5.50% word error rate in streaming mode and 5.04% in non-streaming use — with Google noting multilingual performance has improved over Chirp 3.

Official chart of Gemini 3.5 Transcribe performance on the multilingual FLEURS benchmark
Gemini 3.5 Transcribe on the FLEURS benchmark — official image from Google's blog

On languages specifically: Google's post says the new transcription capabilities detect specialized jargon across more than 85 languages, and the FLEURS benchmark itself covers a wide language set. However, Google has not yet published a detailed per-language availability list for the consumer rollout — the macOS launch is English, and Android's Rambler arrives in unnamed "select countries and languages." If you work in a language other than English, treat availability as unconfirmed until you see it live in your own app.

Where developers and creators find it today

Beyond the API, Google has wired the model into its own surfaces: Gboard, the Antigravity developer platform, the Gemini app — where macOS users can analyze files, generate images, and search using voice alone — and Chrome, which is coming soon. Google also names Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents as platforms building voice-driven interfaces on top of the Gemini Live API, which means transcription and captioning tools powered by this model will keep appearing in third-party products over the coming months.

Official early customer feedback about Gemini 3.5 Transcribe from Google's announcement
Early company reactions — official image from Google's blog

What this means for you as a content creator

If you produce podcasts, videos, or interviews, this class of model changes the economics of your workflow in three concrete ways:

  1. Editing time collapses. The gap between a raw transcript and a publishable one is typically one to three hours per long episode. Automatic filler removal and speaker separation eliminate most of that.
  2. Short-form clips get cheaper. Word-level timestamps turn clip extraction from re-listening into text search. Find the quote, read the timestamp, cut the segment.
  3. Downstream output improves. Translations, show notes, titles, and summaries are all generated from text — the cleaner the source text, the better everything built on top of it.

For small businesses, the same engine applies to customer-call recordings, meetings, and long voice notes: turn them into searchable, archivable summaries and task lists. The public API means you can build exactly that into your existing stack.

Quick comparison with what you probably use today

ToolFiller removalSpeaker labelsTimestampsAccess
Gemini 3.5 TranscribeYesUp to 3 speakersPer wordAPI + Gemini app (gradual rollout)
Google Docs / keyboard dictationNoNoNoFree for everyone
Open-source WhisperNo (literal text)LimitedPer segmentRequires technical setup
Human transcription serviceOn requestYesPossibleCost per minute

The practical takeaway: if your workflow is "talk first, edit later," this generation of transcription models invites you to flip the pipeline — hand the model raw audio and receive an editable draft rather than an exhausting literal copy. When that draft is ready, tools like ARWriter's article and research writer can carry it the rest of the way, from cleaned transcript to structured, publishable article.

Honest limits you should know before relying on it

  • The consumer launch is English-first. macOS availability is explicitly English at launch, and Android's country list is unnamed — do not assume immediate availability in your language until you test it.
  • Three-speaker cap. A four-person panel podcast will still need manual cleanup for the fourth voice.
  • "Polished" cuts both ways. Disfluency removal is great for content but unacceptable when you need a verbatim quote — always check important quotes against the original recording.
  • API pricing was not detailed in the launch post. Check Google's official pricing page before building a service on it.
  • Benchmarks are benchmarks. The FLEURS and Artificial Analysis figures describe the model's test performance, not a guarantee on your specific audio. Run a sample of your real recordings first.

Frequently asked questions

What is Gemini 3.5 Transcribe?

A new Google speech-to-text model announced on August 26, 2026. It converts raw audio into polished, formatted text, removes filler words, attributes up to three speakers, supports custom vocabulary, and handles live language switching.

Is Gemini 3.5 Transcribe available in my language?

Google says the new capabilities cover more than 85 languages for jargon detection and multilingual transcription, but no detailed per-language list has been published. The consumer rollout started in English on macOS and in select countries for Android's Rambler feature.

Is Gemini 3.5 Transcribe free?

Developers access it via the Gemini API in Google AI Studio as a public preview under standard API pricing. Consumers get it gradually inside the Gemini app and Rambler; Google has not announced a separate paid tier for it.

How does it help podcasters and video creators specifically?

It shortens the path from recording to ready text: episodes arrive cleaned, speaker-separated, and word-stamped, which makes clipping, translating, summarizing, and article-building dramatically faster.

How is it different from Chirp 3?

According to Google: better word error rates on FLEURS, a 70% improvement in time-to-final-transcription per Artificial Analysis, plus new capabilities like disfluency removal, speaker attribution, and smart formatting that Chirp 3 did not offer.

Bottom line: a launch like this does not end transcription as a craft, but it moves the competition from "who types the words" to "who understands the audio." Test it on your own recordings when it reaches you, measure the time you save, then plug the output directly into your editing and publishing stack — ARWriter is built to take cleaned text exactly like this and turn it into finished, publishable content. For broader context on Google's recent direction, see our coverage of Gemini 3.7 Flash launching at half the price.