Google ships Gemini 3.8 Live voice models that reason and call tools while talking
Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, its most advanced live dialogue models to date. Both are native speech-to-speech models for real-time voice agents, live in the Gemini Live API and Google AI Studio. The headline trick: they reason and call tools in the background while the conversation continues, instead of going silent while the agent works.
The pair splits by job. 3.8 Live is the scale tier for low-latency conversation; Live Extended Thinking is built for multi-step work like technical diagnosis or tutoring, and it tops Artificial Analysis' Speech-to-Speech Quality Index at 82.6. Google reports 97 language support, near-real-time visual understanding, and audio watermarked with SynthID.
Early enterprise partners include Salesforce, Genspark and Lumeris. The release extends the Gemini push into voice, a front where OpenAI's Realtime API and newer agentic voice startups have been competing — and where the ability to act mid-conversation is becoming the key differentiator.
UPDATE 2026-09-23: the rollout keeps widening. Google’s conversational Search Live now runs on 3.8 Live for more natural voice search, and the Gemini Omni Flash Preview endpoint is deprecated September 30 — developers still on the preview must migrate to 3.8 Live by then. Low-latency audio-to-audio voice models are now Google’s default voice stack.
UPDATE 2026-09-23: migration deadline — Google's earlier Gemini Omni Flash Preview endpoint is scheduled for deprecation on September 30, 2026. Developers still on the preview must move production voice workloads to Gemini 3.8 Live or 3.8 Live Extended Thinking before the cutoff. (ts2.tech, citing Google's blog and Gemini API release notes)