Google’s September 23, 2026 Gemini 3.8 text-to-speech post frames Flash TTS and Flash-Lite TTS as a shift from static presets to a promptable vocal studio — with the same API schema across both models so Alberta operators can swap SKUs on one parameter.
What Google published for 3.8 Flash TTS
Per Google, Gemini 3.8 Flash TTS is built for deep creative direction and character design — natural-language prompts for role, accent, and voice characteristics. Flash-Lite is the high-volume twin for dubbing, audio content, and expressive voice agents. Both share the Gemini 3.8 TTS schema described on the model cards: treat transcript text as verbatim; put sustained delivery in structured speech_metadata; keep angle-bracket tags for point-in-time vocal events.
The blog puts the library at 2,000+ production-ready voices with regional varieties including Mexican Spanish, Quebec French, and Scots English. Voice replication recreates a vocal profile from about a 30-second sample with built-in consent verification. Every Gemini Audio clip is watermarked with SynthID; C2PA credentials are stated alongside. Model cards list input 8,192 and output 16,384 token limits for both SKUs.
“Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio.”Google — Gemini 3.8 text-to-speech says hello, September 23, 2026
Benchmarks and limits (from the primaries only)
Google reports Gemini 3.8 Flash TTS at #1 overall on Hume AI’s Voice Design Benchmark (71.4) and leading accent modeling (60.8). Flash and Flash-Lite take #1 and #2 on Hume’s Overall Quality Index. Blind Voice Arena preference wins are listed for Japanese, Brazilian Portuguese, Vietnamese, MSA, Mexican Spanish, and Hindi — attribute those placements to Google’s post only.
| Line | Flash TTS | Flash-Lite TTS | Notes |
|---|---|---|---|
| API id | gemini-3.8-flash-tts | gemini-3.8-flash-lite-tts | Same schema |
| Languages | 130 | 101 | Model cards |
| Tokens in / out | 8,192 / 16,384 | 8,192 / 16,384 | Both cards |
| Hume Voice Design | 71.4 (#1) | — | Google-reported |
| Replication geo | Not in IL, TX, EEA, UK, CH, India | AI Studio replication | |
This desk does not invent token prices. If your invoice shows CAD, quote that separately from any future Google price card — not from this briefing.
What Alberta operators should do
For voice agents Alberta, bilingual Canada chatbot flows, and chatbot lead-gen, the operator move is to prototype the agent cascade on Flash-Lite (latency / throughput) and reserve Flash for regional accent work — including Quebec French when the script needs it. Keep consent recordings for any 30-second replication path, and treat SynthID + C2PA as the vendor’s transparency story for generated audio.
Pair this briefing with Opcelerate’s AI consulting Alberta path and private AI security lane, and keep reading the city desk at The Super Intelligence Times. Foreign hosted TTS ≠ private Alberta inference when prompts carry client recordings or regulated scripts.
The guardrail
Expressive TTS is not permission to skip consent or geo checks. Replication via AI Studio is unavailable in Illinois, Texas, EEA, UK, Switzerland, and India. Hume scores and Voice Arena placements are Google-reported. Do not treat Gemini 3.8 TTS as AGI, and do not invent list prices for board decks.
gemini-3.8-flash-tts / gemini-3.8-flash-lite-tts, 8,192 / 16,384 tokens, SynthID + C2PA, ~30s replication with consent. Attribute Hume 71.4 / 60.8 and library counts to Google only. No invented $. Route client audio through a private AI security review.