Not established from the available sources.
GOOGLE (DEEPMIND) · MODEL RELEASE TRACKER
Gemini 3.8 Flash TTS
Gemini 3.8 Flash TTS is a text-to-speech model announced by Google in September 2026. It is designed for creative voice design and expressive audio generation — creating bespoke voices from natural-language prompts, directing line-level acting cues, supporting many languages and dialogue features, and integrating across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.CURRENT SNAPSHOT1/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
Not established from the available sources.
Not established from the available sources.
RELEASE TIMELINE
Published, source-backed release events only.
Google DeepMind unveils Gemini 4 Argon, a frontier model with 1M-token context
Google DeepMind announced Gemini 4 Argon, a new frontier multimodal model optimized for long-horizon reasoning and complex workflows. Argon is rolling out initially to trusted cyber defenders via the Fairwind Program, expands context length to 1 million tokens, reports leading benchmark performance across coding, finance, legal, and video understanding, and will be made more widely available after phased safety testing and engagement with U.S. government pre-release processes; Google also published introductory pricing.
Google releases Gemini 3.8 Flash and Flash‑Lite text‑to‑speech models
Google introduced two new Gemini 3.8 text‑to‑speech models—Flash TTS for highly expressive, character-driven voice design and Flash‑Lite TTS for cost‑efficient, high‑volume use—available across Google AI Studio, Gemini API/Enterprise, Gemini Notebook, and Google Vids. The models claim large voice coverage (2,000+ production voices, 100+ languages/dialects), generative voice creation from prompts, 30‑second voice replication with consent verification, SynthID watermarking and C2PA credentials, and features for long‑form and multi‑speaker scene staging.
VERIFIABLE FACTS
Every value stays attached to a source and date.
Announced by Google on 2026-09-23.
Text-to-speech (audio generation) model.
Can create entirely new voices from natural-language prompts with fine-grained control over role, accent, acting cues, pacing, dialect shifts, and backchanneling.
Supports more than 100 languages and dialects.
Supports stage directions per line, two-voice dialogue, and nonverbal sounds (e.g., laughter, sighs).
Has a voice-cloning feature that can build a voice profile from a 30-second audio sample.
Offers a large preset library (reported >2,000 preset voices) and the ability to scale from a limited set to an effectively unlimited/infinite set of original voices.
Google states the model can generate hours of audio with minimal 'speaker drift' (the voice barely changes over time).
WHAT CHANGED