Gemini 3.8 Live Reaches General Availability
Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced real-time voice dialogue models yet, rolling out now across the Gemini API, Google AI Studio, and Gemini Enterprise. The base model targets low-latency, cost-efficient voice agents with visual grounding and background tool execution, while the Extended Thinking variant adds multi-step reasoning and narrates its thought process aloud during complex tasks. Both models automatically detect and switch between 97 languages mid-conversation, and Extended Thinking topped Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. All AI-generated audio carries SynthID watermarking for provenance detection.
Key Takeaways
- Google shipped two new audio-to-audio models simultaneously, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, positioning them as its most advanced conversational voice AI yet.
- Extended Thinking claimed the #1 spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, ahead of competing voice models.
- Both models support automatic detection and switching across 97 languages mid-conversation, removing manual language configuration.
- Extended Thinking introduces simultaneous reasoning and speech, narrating background task progress aloud instead of pausing silently.
- The standard Gemini 3.8 Live model is reported to undercut rivals like GPT-Live-1 Astra on price while still scoring 76.0% on the Speech to Speech Index.
- Every generated audio clip carries SynthID watermarking, Google's provenance signal for detecting AI-generated speech.
Sources & Mentions
4 external resources covering this update
Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
Unite.AI
Google Announces Gemini 3.8 Live and 3.8 Live Extended Thinking
Thurrott.com
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
MarkTechPost
Google Releases Gemini 3.8 Live-Extended Conversational Model, Claims Better Performance Than Rivals At Lower Price
OfficeChai
Two New Voice Models for Real-Time Conversation
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, describing the pair as its most advanced live dialogue models to date. Both are audio-to-audio models built for natural, low-latency conversation, and both are available immediately through the Gemini API and Google AI Studio, with staged rollout to Gemini Enterprise, Search Live, Gemini Live, and Google Workspace.
Gemini 3.8 Live: Built for Scale
Gemini 3.8 Live is positioned as the default choice for most low-latency voice agent experiences. It combines conversational intelligence with fluid dialogue and visual grounding, processing visual inputs in near real time while executing tool calls and API requests in the background without breaking the flow of conversation. Google emphasizes that the model is built for scale and cost efficiency, making it suited to high-volume production voice agents rather than only experimental use.
Gemini 3.8 Live Extended Thinking: Reasoning Out Loud
The Extended Thinking variant is aimed at higher-complexity tasks that require multi-step reasoning. Rather than pausing to think silently, the model reasons and speaks simultaneously, using natural verbal cues such as "Let me check that…" to acknowledge a request while it works. It also provides live progress narration during background task execution, which Google positions as useful for enterprise-grade workflows like multi-step booking coordination or business plan generation.
Language Coverage and Use Cases
Both models support 97 languages with automatic mid-conversation detection and switching, removing the need to manually configure a session's language ahead of time. Google demonstrated the models across a range of use cases, including real-time employee onboarding guidance, visually grounded game play, converting hand-drawn sketches into working React components by voice, and multi-step task coordination that spans several backend calls.
Benchmark Performance
Gemini 3.8 Live Extended Thinking took the top overall position on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, and led agentic task completion benchmarks with 68.6% on Ï„-Voice and 35.1% on Sierra's Ï„-Voice-banking benchmark, alongside a 97.7% score on Big Bench Audio. The standard Gemini 3.8 Live model scored 76.0% on the same Speech to Speech Index while remaining substantially cheaper to run than some competing voice models, according to third-party coverage of the launch.
Safety and Availability
Every piece of AI-generated audio from both models is embedded with SynthID watermarking, an imperceptible signal designed to keep AI-generated speech detectable even after editing or compression. Developers can start building with both models today through the Gemini API and AI Studio, with enterprise and consumer surfaces (Gemini Enterprise, Workspace, Search Live) following in the coming weeks.