Google Debuts Gemini 3.8 Live with Real-Time Reasoning
Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, introducing native voice models that combine real-time reasoning, visual grounding, and background tool execution.

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two native voice models designed to handle streaming audio and visual inputs simultaneously. Available through the Gemini Live API and Google AI Studio, these models allow developers to build voice agents that can maintain spoken dialogue, process camera feeds or shared screens, and execute background tool calls without interrupting the conversation. The models can also automatically detect and switch between 97 supported languages mid-session.
The new models have achieved top marks on several industry benchmarks. Gemini 3.8 Live Extended Thinking secured the top spot on the Artificial Analysis Speech-to-Speech Quality Index with a score of 82.6. It also scored 68.6% on tau-Voice agentic tasks and 97.7% on Big Bench Audio. Meanwhile, the standard Gemini 3.8 Live model placed second in the Speech Agent Arena, and both models reached the Pareto frontier on ServiceNow's EVA-Bench, demonstrating a strong balance between workflow completion and conversational quality.
For developers, the release offers two distinct workload profiles. The standard model is optimized for high-volume, latency-sensitive tasks like customer support, while the Extended Thinking version provides multi-step reasoning and spoken progress updates for complex tasks like debugging or booking coordination. This allows the model to verbally update users on its progress while keeping its internal reasoning private. Integrations are already supported by platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents.
While the models streamline voice architectures by replacing multi-stage pipelines with a single native session, Google has not yet published specific end-to-end latency percentiles, context limits, or pricing details. Developers must also manage the engineering complexities of concurrent tool calls, such as handling user interruptions or stale API results. To address safety, Google applies its SynthID watermark to all generated audio to help identify AI-generated content.
This is our own summary of reporting by AlphaSignal



