On September 15, 2026, Google DeepMind officially launched two new speech and audio models named Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both models were made available immediately through the Gemini Live API as well as the Google AI Studio developer environment. With this release, Google intends to provide software engineers and enterprises with responsive voice systems built specifically for interactive, real-time deployments. The debut highlights an ongoing shift across the industry, moving away from simple text generation toward fully conversational audio agents.
At the technical core of the release is a native speech-to-speech architecture that processes acoustic signals directly end-to-end. Traditional voice assistant pipelines typically relied on transcribing incoming spoken audio into text, analyzing that text semantically, and running the generated output through a separate speech synthesis engine. That multistep approach inevitably introduced noticeable latency and stripped away expressive vocal nuances. Gemini 3.8 Live eliminates these intermediate translation barriers by capturing, evaluating, and outputting spoken sound within a unified modal framework.
A standout feature of the launch is Gemini 3.8 Live Extended Thinking, an edition engineered with enhanced logical reasoning capabilities. This model performs parallel background reasoning while the verbal interaction with the human user continues without interruption. Whereas previous reasoning models required extended processing pauses before replying, Google now executes complex thought operations concurrently in background threads. Users can continue conversing normally while the underlying model evaluates contextual data and structures complex conclusions.
Google DeepMind also highlighted asynchronous tool calling as a major operational enhancement. In conventional agent architectures, invoking external functions often led to call drops or awkward conversational freezes while the system waited for third-party endpoints to return data. Gemini 3.8 Live executes tool calls asynchronously during active sessions, preventing audio dropouts or conversational pauses. This capability allows developers to link models to live databases, web services, and external workflows while maintaining natural, uninterrupted dialogue.
The arrival of the new models triggered immediate adoption across real-time infrastructure and streaming providers. Platforms including LiveKit, Agora, and Fishjam simultaneously announced dedicated integrations for Gemini 3.8 Live. These platforms supply the specialized networking infrastructure and transport protocols necessary for bidirectional audio streaming at minimal latency. Their rapid platform support lowers technical hurdles, enabling development teams to implement responsive voice agents without constructing proprietary streaming networks.
The entertainment industry is emerging as an immediate proving ground for these low-latency audio capabilities. Production teams and digital media publishers are preparing interactive audio dramas where listeners directly influence plot trajectories through natural conversation. The technology is also being targeted at dynamic dubbing workflows, enabling fast and flexible multilingual adaptations of spoken media. Furthermore, broadcasters are testing AI-assisted radio co-hosts that can converse live on air without perceptible delay alongside human presenters.

