Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RELEASE · MODELS · #875

Google releases Gemini 3.8 Flash and Flash‑Lite text‑to‑speech models

Google introduced two new Gemini 3.8 text‑to‑speech models—Flash TTS for highly expressive, character-driven voice design and Flash‑Lite TTS for cost‑efficient, high‑volume use—available across Google AI Studio, Gemini API/Enterprise, Gemini Notebook, and Google Vids. The models claim large voice coverage (2,000+ production voices, 100+ languages/dialects), generative voice creation from prompts, 30‑second voice replication with consent verification, SynthID watermarking and C2PA credentials, and features for long‑form and multi‑speaker scene staging.

KEY POINTS

  1. Google introduced two new Gemini 3.8 text‑to‑speech models—Flash TTS for highly expressive, character-driven voice design and Flash‑Lite TTS for cost‑efficient, high‑volume use—available across Google AI Studio, Gemini API/Enterprise, Gemini Notebook, and Google Vids.
  2. The models claim large voice coverage (2,000+ production voices, 100+ languages/dialects), generative voice creation from prompts, 30‑second voice replication with consent verification, SynthID watermarking and C2PA credentials, and features for long‑form and multi‑speaker scene staging.
  3. This release expands generative audio capabilities for creators and enterprises—enabling bespoke, long‑form and multi‑speaker voice production at scale while highlighting voice‑replication and provenance safeguards.
MERIDIAN INTELLIGENCE

DECISION BRIEF

75/100
CONFIDENCE
01

WHAT CHANGED

Google announced two new Gemini 3.8 text-to-speech models: Gemini 3.8 Flash TTS (expressive, character-driven) and Gemini 3.8 Flash‑Lite TTS (cost‑efficient, high‑volume).

02

WHY NOW

The release expands Gemini’s audio capabilities with generative voice design, a large preset voice library, multi‑speaker and long‑form staging features, and a stated focus on scaling expressive audio for creators, developers, and enterprises.

03

WHO IS AFFECTED

Directly affected: creators and creative teams (games, audiobooks, podcasts), developers building voice-enabled products, dubbing and high-volume audio teams, and enterprises using Gemini products (Gemini API, Google AI Studio, Gemini Notebook, Google Vids per vendor messaging).

04

CONFIRMED

All items below are explicitly stated in the supplied source excerpts: - Google introduced two Gemini 3.8 TTS models named Flash TTS and Flash‑Lite TTS (IDs: 1035, 1036, 1048). - Flash TTS is described as for deep creative direction and character design; Flash‑Lite TTS is described as optimized for high-volume, cost-efficient scale (IDs: 1035, 1036, 1048). - The models support generative voice creation from text prompts and control over role, accent, pacing, and expressive nuance across more than 100 languages/dialects (IDs: 1035, 1036, 1048). - Google cites an expansive library of 2,000+ production-ready voices and regional variants like Mexican Spanish, Quebec French, and Scots English (IDs: 1035, 1036, 1048). - The models include features for staging multi‑speaker scenes, line‑by‑line stage directions, and nonverbal sounds (IDs: 1035, 1036, 1048). - A voice cloning/replication feature is described that can build a voice profile from a ~30‑second audio sample and requires a recorded statement of consent that matches the sample (ID: 1048). - Google/DeepMind list availability through Google AI Studio, Gemini API, Gemini Notebook, and Google Vids in vendor posts (IDs: 1035, 1036).

05

UNCERTAIN

Missing or conflicting evidence (explicitly flagged): - Conflicting claims on immediate enterprise/API availability: vendor posts (IDs: 1035, 1036) state availability across Gemini API and Gemini Enterprise, while The Decoder (ID: 1048) reports roll‑out via Gemini API and Google AI Studio and says Gemini Enterprise API access will follow. The supplied excerpts do not reconcile this difference. - C2PA credentials: the supplied excerpts explicitly mention an inaudible SynthID watermark in generated clips (ID: 1048) but do not provide explicit, excerpted confirmation that C2PA credentials are used; the presence/status of C2PA in these models is not shown in the provided excerpts. - Operational details not present in excerpts: pricing, rate limits, latency, SDKs/SDK updates, precise consent verification workflow, abuse‑mitigation controls beyond SynthID, and measured robustness of long‑form generation and speaker‑drift claims are not documented in the provided excerpts. - Real-world quality/coverage across all claimed languages and dialects, and empirical performance of the 30‑second cloning under varied audio conditions, are not evidenced in the supplied text.

06

WHAT TO WATCH

Concrete observable signals to watch next (where to look in vendor/press channels): - Official Gemini product pages, Gemini API / Gemini Enterprise announcements for clarifying enterprise/API availability and rollout timing (IDs: 1035, 1036, 1048). - Google AI Studio / Gemini Notebook / Google Vids documentation and change logs for integration details, sample demos, and supported workflows (IDs: 1035, 1036). - Release notes or blog updates mentioning Voice Remixing availability and any changes to the 30‑second voice cloning workflow or consent verification (ID: 1048). - Technical docs or security/privacy pages describing SynthID implementation and any mention of C2PA credentials or provenance tooling (ID: 1048). - Independent demos, technical evaluations, or media reports demonstrating long‑form stability, speaker‑drift metrics, and cross‑language quality (corroborating vendor claims in IDs: 1035, 1036, 1048).

WHY IT MATTERS

This release expands generative audio capabilities for creators and enterprises—enabling bespoke, long‑form and multi‑speaker voice production at scale while highlighting voice‑replication and provenance safeguards.

EVIDENCE MAP

4

Editorial claims linked to specific sources, with support, contradiction and context shown separately.

Flash TTS is described as aimed at expressive, character‑driven voice design; Flash‑Lite TTS is described as aimed at high‑volume, cost‑efficient use.

SUPPORTED

Checked 2026-09-24 · 3 supporting

Google's posts state the models are available across Google AI Studio, Gemini API/Enterprise, Gemini Notebook, and Google Vids.

SUPPORTED

Checked 2026-09-24 · 2 supporting

The models are described as supporting more than 100 languages and dialects and enabling generative voice creation from text prompts.

SUPPORTED

Checked 2026-09-24 · 3 supporting

SOURCES & TIMELINE

3
01
GOOGLE FIRST-PARTY
Gemini 3.8 text-to-speech says hello

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Director, Research Science, on Behalf of the Gemini Audio Team Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice genera…

↗
02
GOOGLE DEEPMIND FIRST-PARTY
Gemini 3.8 text-to-speech says hello

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Director, Research Science, on Behalf of the Gemini Audio Team Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice genera…

↗
03
GOOGLE FIRST-PARTY
Introducing Gemini 3.8 Live with Live Avatar

Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live Avatar enables a more natural and intuitive conversational experience for enterprises and their users. Software Engineer, on behalf of the Gemini Audio Team Building on the momentum of last week's Gemini 3.8 Live launch, today we …

↗
04
GOOGLE DEEPMIND FIRST-PARTY
Introducing Gemini 3.8 Live with Live Avatar

Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live Avatar enables a more natural and intuitive conversational experience for enterprises and their users. Software Engineer, on behalf of the Gemini Audio Team Building on the momentum of last week's Gemini 3.8 Live launch, today we …

↗
05
THE DECODER INDEPENDENT COVERAGE
Google's new Flash TTS models let you design AI voices from scratch using text descriptions

Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and a voice cloning feature builds voice profiles from 30-second audio samples. Flash TTS is aimed at creative uses such as podcasts, audiobooks, and game characters, while Flash-Lite TTS is designed for low-cost speech generation a…

↗