Google Just Turned Text-to-Speech Into a Voice Studio — Here Is Everything That Actually Shipped
On September 22, 2026, Google made its next-generation speech models generally available, and followed up the next day with a detailed announcement on its official blog titled "Gemini 3.8 text-to-speech says hello." Two models arrived at once: Gemini 3.8 Flash TTS, the flagship built for creative direction and character-driven audio, and Gemini 3.8 Flash-Lite TTS, the cost-efficient workhorse aimed at high-volume dubbing and always-on voice agents. If you produce podcasts, audiobooks, dubbed video, or social clips, the headline is simple: generated speech has crossed from "readable" to "directable," and Google published benchmark results — including blind human preference wins in Modern Standard Arabic — to back the claim.
This guide assembles the full picture directly from Google's primary sources: what genuinely shipped, what an hour of generated audio really costs, where you can try the models for free today, and the honest limitations to weigh before you rebuild your production pipeline around them.

What Google Actually Launched
This is not one model but a coordinated audio system with four connected pieces:
- Gemini 3.8 Flash TTS (
gemini-3.8-flash-tts): the flagship creative model, engineered for studio-grade voice fidelity, nuanced acting, regional dialects, and long-form multi-turn stability. - Gemini 3.8 Flash-Lite TTS (
gemini-3.8-flash-lite-tts): the fast, cost-efficient sibling built to replacegemini-3.1-flash-tts-previewfor high-throughput production and real-time voice agent cascades. - A new Voices endpoint (
/v1beta/voices): an API surface for querying and managing more than 150 prebuilt and custom voices programmatically. - Three companion capabilities: generative voice design, consent-verified voice replication, and the expanded voice library.
Both models are live now for developers through the Gemini API and Google AI Studio, for everyone inside Gemini Notebook (the full Flash model) and Google Vids (the Lite model), with API access through Gemini Enterprise described as "coming soon" in Google's own words. The release rounds out a Gemini Audio family that already includes Gemini 3.8 Live for real-time voice conversation, 3.5 Transcribe for transcription, and 3.5 Live Translate for instant translation.
The Quality Claims, With the Receipts
Google did not ship this release on adjectives alone. On Hume AI's Voice Design Benchmark, the flagship model took the overall #1 spot with a score of 71.4, and also led accent modeling at 60.8. The two new models hold #1 and #2 on Hume AI's Overall Quality Index, with Google reporting major gains over Gemini 3.1 Flash TTS specifically in long-form content and dual-speaker screenplay control.
More interesting for international creators are the blind human preference runs on Voice Arena: listeners rated the new models at "top positions amongst competitors" in key global languages, explicitly including Modern Standard Arabic, alongside Japanese, Brazilian Portuguese, Vietnamese, Mexican Spanish, and Hindi. Google also claims support for over 100 languages in voice design, and its official documentation lists both Standard Arabic and Egyptian Arabic as supported for speech generation. For anyone who has winced through machine-read Arabic in years past, that is the difference between "supported" and "credible."
These two official charts from the announcement summarize where the models landed:


From 30 Voices to More Than 2,000
Before this release, developers picked from roughly 30 fixed preset voices. Google now talks about more than 2,000 production-ready voices with broad language coverage that reaches down to regional varieties — Mexican Spanish, Quebec French, and Scots English are the examples Google itself calls out — signaling sub-dialect granularity rather than just major-language checkboxes.
And because no stock library satisfies a brand that wants a recognizable voice of its own, three tools let you build custom voices:
- Generative voice design: describe a voice in plain language — "a calm Gulf-region female narrator with a warm tone for children's stories" — and the model generates the full vocal persona across role, accent, and characteristics, in more than 100 languages and dialects. Google's own demo examples include a high-energy DJ from Melbourne, a super-tinny monotone robot, and a Japanese dragon brought to life.
- Voice replication: a 30-second recording of your own voice, or one you have the rights to use, is enough to recreate a consistent vocal profile. The system enforces consent: a verbal consent recording from the voice owner must match the reference speaker before any cloned voice is created.
- Save and scale: designed voices persist, so the same identity stays consistent across episodes and projects with minimal drift, instead of drifting session by session.
A fourth capability, Voice Remixing, is announced as coming soon: take an existing library voice and fine-tune its timbre, pitch, pace, and accent with prompts like "add a subtle Southern US accent" or "soften the delivery."
Direct the Performance, Line by Line
This is the structural change that separates this generation from every TTS tool you have already tried. The models do not receive "text" and return "a reading." They accept a written performance:
- Line-by-line direction: write your own stage directions, or let Gemini infer delivery cues from the script itself — from a calm customer-service agent to a whispered suspense scene.
- Native two-speaker scene staging: a full dialogue between two distinct, separated voices with natural conversational turn-taking, generated from a single script. No more recording each side separately and stitching in an editor.
- Realistic conversational texture: explicit tags for non-verbal bursts such as
<laughs>,<sigh>, and<gasp>, plus active-listening interjections like a soft "mhm" or "yeah" placed exactly where they land — comedy timing and reaction beats become writable. - Long-form stability: voice quality, natural pacing, and character timbre hold across hours of continuous audio with minimal speaker drift — the threshold that makes full audiobook generation from a single script realistic.
In practice, a podcaster can now write an episode as a two-character script, receive it fully performed, then revise the delivery of one line without regenerating the rest. That philosophy — script first, then distribution across formats — is exactly how we built the Auto-Writer inside ARWriter, where a well-structured draft flows out to every channel you publish on.
Pricing in Real Numbers: What an Hour of Audio Costs
Google's pricing page lists per-million-token rates, and the key conversion is that one second of audio equals 25 tokens. Here is the full picture:
| Line item | Gemini 3.8 Flash TTS | Gemini 3.8 Flash-Lite TTS | Previous 3.1 Flash TTS |
|---|---|---|---|
| Text input (per 1M tokens) | $0.50 through Dec 31, 2026 (then $1.00) | $0.50 through Dec 31, 2026 (then $1.00) | $1.00 |
| Audio output (per 1M tokens) | $9.00 through Dec 31, 2026 (then $18.00) | $6.00 through Dec 31, 2026 | $20.00 |
| Free tier | Available | Available | Available |
Translating into product language: one million audio tokens equals 40,000 seconds, or about 11.1 hours of audio. That means the flagship model at $9 per million works out to roughly $0.81 per generated hour, and the Lite model to about $0.54 per hour — at the promotional rates running through December 31, 2026. Against the previous generation's $20 per million audio tokens, the new flagship is more than 55% cheaper while benchmarking dramatically higher.
One planning caveat that matters: these are launch prices. On January 1, 2027, Flash audio output doubles to $18 per million tokens (Lite's post-promo rate follows the same pattern per Google's pricing page). Build your unit economics with the 2027 numbers and treat 2026 as a grace period, not the baseline.
Where to Try It Today: A Practical Walkthrough
The fastest entry point is the new audio playground in Google AI Studio, which Google describes as a voice design workspace: describe a voice, hear it, save it, then move it into a dual-speaker screenplay editor and direct the delivery line by line. Voice replication of your own voice is included in the playground experience.
- Open aistudio.google.com and find the speech generation section.
- Describe the vocal persona in natural language — performance style, approximate age, tone, accent — and iterate until it fits.
- Paste a passage from your own script and experiment with inline directions such as "building excitement" or "calm newsreader delivery," then compare takes.
- Developers can call
gemini-3.8-flash-ttsorgemini-3.8-flash-lite-ttsdirectly in one API call, with voice management through/v1beta/voices.
Outside the developer path, the full Flash model is rolling out inside Gemini Notebook, and the Lite model inside Google Vids for narration over your slides and clips without leaving Google's tools. Teams that need enterprise compliance features should note that Gemini Enterprise availability is explicitly "coming soon" rather than live.
For a complete short-video or article workflow — idea, script, structure, publishing — the practical pattern is to draft in a writing tool you trust (our own is ARWriter, built Arabic-first) and then pass the finished script to whichever audio engine wins your ear test.

Quick Comparison: Where This Sits Against Your Current Options
The reflex is to ask "is this better than ElevenLabs?" — but any article that declares a single winner without knowing your use case is guessing. Three factors decide it for most workflows. First, documented Arabic and multilingual quality: this is where Google's Voice Arena results, explicitly covering Modern Standard Arabic, carry weight. Second, cost at volume: at roughly half a dollar to a dollar per finished hour at launch rates, the economics are competitive with anything in the market. Third, ecosystem fit: if you already live in Google's workspace — Vids, Notebook, AI Studio's free tier for testing — the integration path is frictionless. Creators who need a specific licensed actor's voice or deep libraries of niche regional accents will still want to A/B their actual scripts across two or three engines before committing; the only test that matters is your own ears on your own text.
Inside Google's own lineup the boundaries are cleaner than ever: this release is the "recording studio," Gemini 3.8 Live is the "live conversation," and dedicated transcription tools fill the third role. We compared the transcription side in depth in our coverage of OpenAI's GPT-Live-1 voice API earlier this month. Three tools, three distinct jobs.
Honest Limitations to Know Before You Commit
- Promotional pricing expires: launch rates end December 31, 2026, and Flash output doubles on January 1, 2027 — model your margins on the higher number.
- Regional restrictions on replication: voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland, or India. The Gulf and wider MENA region face no listed restriction — worth documenting for our readers.
- Permanent watermarking: every generated clip carries Google's SynthID watermark woven into the audio waveform, plus C2PA credentials for replicated voices. Excellent for transparency and disclosure obligations; it also means impersonation use cases are architecturally closed off — by design.
- Voice Remixing is not live yet: it is listed as coming soon, so keep it out of near-term production plans.
- Enterprise gap: compliance-grade deployments wait on the Gemini Enterprise API rollout.
What This Means for You: Three Ready Workflows
1) Short-form video creators: write a 45-second script, generate the voiceover with Lite (well under a cent per clip), and lay it over your edit. Professional-grade narration with no microphone and no studio.
2) Audiobook and podcast producers: the flagship model plus two-speaker staging plus long-form stability means a full chapter performed from the chapter text — a narrator and a character voice, both saved and persistent, holding their identities across every future chapter.
3) Dubbing and regional expansion: Google's named launch partners for this release are themselves dubbing and localization companies — HeyGen, Linguana, Wondercraft, and Ollang — which tells you the flagship use case the industry sees: moving content between languages at scale, in both directions, including into and out of Arabic. Anyone sitting on an existing content library is looking at the cheapest dubbing mechanism the market has offered to date.
Notice the pattern across all three: audio is no longer a separate production stage that happens after writing. Voice direction — tone, emotion, pacing — is now written into the script at the moment of composition. Teams that learn to write "speakable text" in this sense, with short sentences, intentional pauses, and embedded delivery notes, will collapse the distance from idea to published audio dramatically.
The Bottom Line
Overnight, Modern Standard Arabic gained a speech model that tops blind human preference rankings, at a cost approaching a dollar per finished hour, inside tools you can test for free this afternoon. Competitors will answer — that is guaranteed — and every response will be good news for creators who work in languages the industry historically treated as an afterthought. Run the free AI Studio playground on your own script today and judge with your ears, not press headlines. When your scripts are ready, you can draft and structure them end to end with ARWriter's writing tools — built for creators who publish across languages, Arabic first.
Official Sources
- Google's official announcement: Gemini 3.8 text-to-speech says hello (Google blog, September 23, 2026)
- Gemini API release notes — general availability of both models (September 22, 2026)
- Official Gemini API pricing page
- Speech generation documentation and language table — includes Standard Arabic and Egyptian Arabic
- Gemini 3.8 Audio model card (Google DeepMind)
Documented event dates: general availability September 22, 2026; announcement post September 23, 2026. This is an analytical guide to a several-days-old release, not breaking coverage.
Frequently Asked Questions
Does Gemini 3.8 Flash TTS support Arabic?
Yes. Google's official documentation lists both Standard Arabic and Egyptian Arabic as supported, and the announcement specifically states the models earned top positions in blind Voice Arena preference ratings in Modern Standard Arabic alongside Japanese, Brazilian Portuguese, Vietnamese, Mexican Spanish, and Hindi.
Is Gemini 3.8 Flash TTS free to use?
A free tier is available through Google AI Studio and the Gemini API within standard free usage limits. Paid pricing starts at $0.50 per million text tokens and $6–$9 per million audio output tokens through December 31, 2026.
How much does it cost to convert an article or book into audio?
Since one second of audio consumes 25 tokens, one million audio tokens equals about 11.1 hours of sound. That works out to roughly $0.54 per hour with Flash-Lite and about $0.81 per hour with the flagship model at launch pricing through the end of 2026.
Can I clone my own voice?
Yes. Voice Replication needs just a 30-second sample plus a verbal consent recording that the system verifies against the reference voice. Through AI Studio it is unavailable in Illinois, Texas, the EEA, the UK, Switzerland, and India; no restriction is listed for the Gulf or the wider MENA region.
What is the difference between Gemini 3.8 Flash TTS and Gemini 3.8 Live?
Flash TTS converts written text into performed audio for recorded content such as podcasts, audiobooks, and dubbing. Gemini 3.8 Live is a live two-way conversational model that listens and responds in real time. Both belong to the Gemini Audio family, but they serve entirely different production jobs.