GOOGLE / GOOGLE DEEPMIND · MODEL RELEASE TRACKER
Gemini 3.8 Live
Gemini 3.8 Live is a live dialogue model in the Gemini family designed for near-real-time speech-first interactions that can process visual inputs, call tools asynchronously, and operate at scale for voice agents. A Live Avatar variant adds low-latency streaming video with lip-sync and expressions and is offered to enterprise customers.CURRENT SNAPSHOT5/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
RELEASE TIMELINE
Published, source-backed release events only.
Google DeepMind unveils Gemini 4 Argon, a frontier model with 1M-token context
Google DeepMind announced Gemini 4 Argon, a new frontier multimodal model optimized for long-horizon reasoning and complex workflows. Argon is rolling out initially to trusted cyber defenders via the Fairwind Program, expands context length to 1 million tokens, reports leading benchmark performance across coding, finance, legal, and video understanding, and will be made more widely available after phased safety testing and engagement with U.S. government pre-release processes; Google also published introductory pricing.
Google releases Gemini 3.8 Flash and Flash‑Lite text‑to‑speech models
Google introduced two new Gemini 3.8 text‑to‑speech models—Flash TTS for highly expressive, character-driven voice design and Flash‑Lite TTS for cost‑efficient, high‑volume use—available across Google AI Studio, Gemini API/Enterprise, Gemini Notebook, and Google Vids. The models claim large voice coverage (2,000+ production voices, 100+ languages/dialects), generative voice creation from prompts, 30‑second voice replication with consent verification, SynthID watermarking and C2PA credentials, and features for long‑form and multi‑speaker scene staging.
VERIFIABLE FACTS
Every value stays attached to a source and date.
131,072 input tokens
65,536 output tokens
Audio with optional text transcription; Live response modality is audio
2026-09-15
$0.75 USD / 1M text input tokens; paid standard Live API
$4.5 USD / 1M text output tokens; paid standard Live API
$3 USD / 1M audio input tokens; paid standard Live API
$12 USD / 1M audio output tokens; paid standard Live API
$1 USD / 1M image/video input tokens; paid standard Live API
Announced as a new live dialogue model on 2026-09-15 (introduced alongside Gemini 3.8 Live Extended Thinking).
Processes visual inputs in near real-time to enrich conversations and provide visual grounding.
Supports asynchronous function calling: it can execute API and tool calls in the background while continuing to stream audio responses to the user.
Near‑real‑time live dialogue (speech‑to‑speech) with support for processing visual inputs (visual grounding) during conversations.
Available to developers via the Live API and Google AI Studio; accessible on the Gemini Enterprise Agent Platform. The Live Avatar feature (near‑real‑time video+speech persona) is available in Gemini Enterprise.
Gemini 3.8 Live is available to developers via the Gemini API / Live API and Google AI Studio.
Listed Live API pricing: $0.005 per minute for audio input and $0.018 per minute for audio output (reported in developer release notes).
Reported to have strong user preference, securing second place in the Speech Agent Arena (as stated in the announcement).
Supports asynchronous/background function and tool calling (execute API/tool calls while continuing audio responses), alphanumeric precision (accurate parsing of codes/numbers), automatic multilingual detection/transition (reported support for ~97 languages), incremental content updates, and maintaining dialogue while performing tasks.
Gemini 3.8 Live showed high user preference, securing second place in the Speech Agent Arena (reported by Google/DeepMind posts about the 3.8 Live launch).
In the VAmoS voice-agent benchmark, Grok Voice led; Gemini 3.8 Live and GPT‑Live followed, with Gemini 3.8 Live reportedly at about the same cost per call as GPT‑Live (VAmoS calibration and evaluation results).
In the VAmoS evaluation, the presence of background television greatly reduced pooled task completion (from 38.7% to 8.6%), indicating that voice-agent performance (including stacks evaluated with Gemini 3.8 Live) can be highly sensitive to competing background speech/noise.
WHAT CHANGED