Tech Meridian← ALL MODELS
PROMY MERIDIAN

GOOGLE / GOOGLE DEEPMIND · MODEL RELEASE TRACKER

Gemini 3.8 Live

Gemini 3.8 Live is a live dialogue model in the Gemini family designed for near-real-time speech-first interactions that can process visual inputs, call tools asynchronously, and operate at scale for voice agents. A Live Avatar variant adds low-latency streaming video with lip-sync and expressions and is offered to enterprise customers.

CURRENT SNAPSHOT5/5 DIMENSIONS WITH DATA

The dimensions that change the decision.

PRICING
Input price — text$0.75 USD / 1M text input tokens; paid standard Live API
Output price — text$4.5 USD / 1M text output tokens; paid standard Live API
Input price — audio$3 USD / 1M audio input tokens; paid standard Live API
Output price — audio$12 USD / 1M audio output tokens; paid standard Live API
Input price — image/video$1 USD / 1M image/video input tokens; paid standard Live API
Pricing (Live API)Listed Live API pricing: $0.005 per minute for audio input and $0.018 per minute for audio output (reported in developer release notes).
CONTEXT WINDOW
Input token limit131,072 input tokens
MODALITIES
Output modalitiesAudio with optional text transcription; Live response modality is audio
Visual input processingProcesses visual inputs in near real-time to enrich conversations and provide visual grounding.
ModalitiesNear‑real‑time live dialogue (speech‑to‑speech) with support for processing visual inputs (visual grounding) during conversations.
BENCHMARKS
User-preference benchmarkGemini 3.8 Live showed high user preference, securing second place in the Speech Agent Arena (reported by Google/DeepMind posts about the 3.8 Live launch).
VAmoS benchmark mentionIn the VAmoS voice-agent benchmark, Grok Voice led; Gemini 3.8 Live and GPT‑Live followed, with Gemini 3.8 Live reportedly at about the same cost per call as GPT‑Live (VAmoS calibration and evaluation results).
AVAILABILITY
AvailabilityAvailable to developers via the Live API and Google AI Studio; accessible on the Gemini Enterprise Agent Platform. The Live Avatar feature (near‑real‑time video+speech persona) is available in Gemini Enterprise.
Developer availabilityGemini 3.8 Live is available to developers via the Gemini API / Live API and Google AI Studio.

RELEASE TIMELINE

Published, source-backed release events only.

5 SOURCES · IMPORTANCE 9.0

Google DeepMind unveils Gemini 4 Argon, a frontier model with 1M-token context

Google DeepMind announced Gemini 4 Argon, a new frontier multimodal model optimized for long-horizon reasoning and complex workflows. Argon is rolling out initially to trusted cyber defenders via the Fairwind Program, expands context length to 1 million tokens, reports leading benchmark performance across coding, finance, legal, and video understanding, and will be made more widely available after phased safety testing and engagement with U.S. government pre-release processes; Google also published introductory pricing.

→
5 SOURCES · IMPORTANCE 8.0

Google releases Gemini 3.8 Flash and Flash‑Lite text‑to‑speech models

Google introduced two new Gemini 3.8 text‑to‑speech models—Flash TTS for highly expressive, character-driven voice design and Flash‑Lite TTS for cost‑efficient, high‑volume use—available across Google AI Studio, Gemini API/Enterprise, Gemini Notebook, and Google Vids. The models claim large voice coverage (2,000+ production voices, 100+ languages/dialects), generative voice creation from prompts, 30‑second voice replication with consent verification, SynthID watermarking and C2PA credentials, and features for long‑form and multi‑speaker scene staging.

→

VERIFIABLE FACTS

Every value stays attached to a source and date.

AVAILABILITY · AvailabilityDEVELOPER CLAIM

Available to developers via the Live API and Google AI Studio; accessible on the Gemini Enterprise Agent Platform. The Live Avatar feature (near‑real‑time video+speech persona) is available in Gemini Enterprise.

CAPABILITIES · Key capabilitiesDEVELOPER CLAIM

Supports asynchronous/background function and tool calling (execute API/tool calls while continuing audio responses), alphanumeric precision (accurate parsing of codes/numbers), automatic multilingual detection/transition (reported support for ~97 languages), incremental content updates, and maintaining dialogue while performing tasks.

LIMITATIONS · Performance sensitivity to background audio (VAmoS)INDEPENDENTLY SUPPORTED

In the VAmoS evaluation, the presence of background television greatly reduced pooled task completion (from 38.7% to 8.6%), indicating that voice-agent performance (including stacks evaluated with Gemini 3.8 Live) can be highly sensitive to competing background speech/noise.

WHAT CHANGED

Stored passport versions, without reconstructed history.

Passport updated21 facts
MODALITIES · Multimodal (audio + low-latency video)− Paired with Live Avatar, couples near-real-time video generation with speech to create an experience that listens, sees, and speaks with a dynamic visual persona (precise lip-syncing and natural expressions).
AVAILABILITY · Availability+ Available to developers via the Live API and Google AI Studio; accessible on the Gemini Enterprise Agent Platform. The Live Avatar feature (near‑real‑time video+speech persona) is available in Gemini Enterprise.
CAPABILITIES · Key capabilities+ Supports asynchronous/background function and tool calling (execute API/tool calls while continuing audio responses), alphanumeric precision (accurate parsing of codes/numbers), automatic multilingual detection/transition (reported support for ~97 languages), incremental content updates, and maintaining dialogue while performing tasks.
PRICING · Pricing (Live API)Listed Live API pricing: $0.005 per minute for audio input and $0.018 per minute for audio output.→Listed Live API pricing: $0.005 per minute for audio input and $0.018 per minute for audio output (reported in developer release notes).
AVAILABILITY · Live Avatar availability− The Live Avatar variant (Gemini 3.8 Live with Live Avatar) is available to Gemini Enterprise customers (stated as available starting today).
MODALITIES · Modalities+ Near‑real‑time live dialogue (speech‑to‑speech) with support for processing visual inputs (visual grounding) during conversations.
LIMITATIONS · Performance sensitivity to background audio (VAmoS)+ In the VAmoS evaluation, the presence of background television greatly reduced pooled task completion (from 38.7% to 8.6%), indicating that voice-agent performance (including stacks evaluated with Gemini 3.8 Live) can be highly sensitive to competing background speech/noise.
BENCHMARKS · User-preference benchmark+ Gemini 3.8 Live showed high user preference, securing second place in the Speech Agent Arena (reported by Google/DeepMind posts about the 3.8 Live launch).
BENCHMARKS · VAmoS benchmark mention+ In the VAmoS voice-agent benchmark, Grok Voice led; Gemini 3.8 Live and GPT‑Live followed, with Gemini 3.8 Live reportedly at about the same cost per call as GPT‑Live (VAmoS calibration and evaluation results).
Passport updated17 facts
AVAILABILITY · Developer availabilityGemini 3.8 Live is available to developers via the Gemini API / Live API and Google AI Studio.→Gemini 3.8 Live is available to developers via the Gemini API / Live API and Google AI Studio.
OTHER · Asynchronous tool / API callingSupports asynchronous function calling: it can execute API and tool calls in the background while continuing to stream audio responses to the user.→Supports asynchronous function calling: it can execute API and tool calls in the background while continuing to stream audio responses to the user.
MODALITIES · Multimodal (audio + low-latency video)Paired with Live Avatar, couples near-real-time video generation with speech to create an experience that listens, sees, and speaks with a dynamic visual persona (precise lip-syncing and natural expressions).→Paired with Live Avatar, couples near-real-time video generation with speech to create an experience that listens, sees, and speaks with a dynamic visual persona (precise lip-syncing and natural expressions).
CONTEXT WINDOW · Input token limit+ 131,072 input tokens
RELEASE · Introduction / releaseMeasurement units or comparison conditions updated
PRICING · Pricing (Live API, audio)Listed Live API pricing: $0.005 per minute for audio input and $0.018 per minute for audio output.→Listed Live API pricing: $0.005 per minute for audio input and $0.018 per minute for audio output.
AVAILABILITY · Live Avatar availabilityThe Live Avatar variant (Gemini 3.8 Live with Live Avatar) is available to Gemini Enterprise customers (stated as available starting today).→The Live Avatar variant (Gemini 3.8 Live with Live Avatar) is available to Gemini Enterprise customers (stated as available starting today).
CAPABILITIES · Maximum output+ 65,536 output tokens
MODALITIES · Output modalities+ Audio with optional text transcription; Live response modality is audio
PRICING · Input price — audio+ $3 USD / 1M audio input tokens; paid standard Live API
PRICING · Input price — image/video+ $1 USD / 1M image/video input tokens; paid standard Live API
PRICING · Input price — text+ $0.75 USD / 1M text input tokens; paid standard Live API
PRICING · Output price — audio+ $12 USD / 1M audio output tokens; paid standard Live API
PRICING · Output price — text+ $4.5 USD / 1M text output tokens; paid standard Live API
RELEASE · Release date+ 2026-09-15
OTHER · User preference / rankingReported to have strong user preference, securing second place in the Speech Agent Arena (as stated in the announcement).→Reported to have strong user preference, securing second place in the Speech Agent Arena (as stated in the announcement).
MODALITIES · Visual input processingProcesses visual inputs in near real-time to enrich conversations and provide visual grounding.→Processes visual inputs in near real-time to enrich conversations and provide visual grounding.
Passport created8 facts