Skip to content
AI ConnectPowered by VELENTIS
AI-generated2 min

Google DeepMind Introduces Gemini 3.8 Live for Speech-to-Speech Applications

Google DeepMind has launched Gemini 3.8 Live, offering native speech-to-speech processing with parallel background reasoning and uninterrupted asynchronous tool calling.

This article was AI-generated and published automatically. Context, labelling and all sources at the end of the article.

(KI-generiertes Symbolbild: Gemini / AI Connect)

On September 15, 2026, Google DeepMind officially launched two new speech and audio models named Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both models were made available immediately through the Gemini Live API as well as the Google AI Studio developer environment. With this release, Google intends to provide software engineers and enterprises with responsive voice systems built specifically for interactive, real-time deployments. The debut highlights an ongoing shift across the industry, moving away from simple text generation toward fully conversational audio agents.

At the technical core of the release is a native speech-to-speech architecture that processes acoustic signals directly end-to-end. Traditional voice assistant pipelines typically relied on transcribing incoming spoken audio into text, analyzing that text semantically, and running the generated output through a separate speech synthesis engine. That multistep approach inevitably introduced noticeable latency and stripped away expressive vocal nuances. Gemini 3.8 Live eliminates these intermediate translation barriers by capturing, evaluating, and outputting spoken sound within a unified modal framework.

A standout feature of the launch is Gemini 3.8 Live Extended Thinking, an edition engineered with enhanced logical reasoning capabilities. This model performs parallel background reasoning while the verbal interaction with the human user continues without interruption. Whereas previous reasoning models required extended processing pauses before replying, Google now executes complex thought operations concurrently in background threads. Users can continue conversing normally while the underlying model evaluates contextual data and structures complex conclusions.

Google DeepMind also highlighted asynchronous tool calling as a major operational enhancement. In conventional agent architectures, invoking external functions often led to call drops or awkward conversational freezes while the system waited for third-party endpoints to return data. Gemini 3.8 Live executes tool calls asynchronously during active sessions, preventing audio dropouts or conversational pauses. This capability allows developers to link models to live databases, web services, and external workflows while maintaining natural, uninterrupted dialogue.

The arrival of the new models triggered immediate adoption across real-time infrastructure and streaming providers. Platforms including LiveKit, Agora, and Fishjam simultaneously announced dedicated integrations for Gemini 3.8 Live. These platforms supply the specialized networking infrastructure and transport protocols necessary for bidirectional audio streaming at minimal latency. Their rapid platform support lowers technical hurdles, enabling development teams to implement responsive voice agents without constructing proprietary streaming networks.

The entertainment industry is emerging as an immediate proving ground for these low-latency audio capabilities. Production teams and digital media publishers are preparing interactive audio dramas where listeners directly influence plot trajectories through natural conversation. The technology is also being targeted at dynamic dubbing workflows, enabling fast and flexible multilingual adaptations of spoken media. Furthermore, broadcasters are testing AI-assisted radio co-hosts that can converse live on air without perceptible delay alongside human presenters.

What this means for you

For developers and media companies, Gemini 3.8 Live marks a departure from cumbersome, multi-step voice pipelines based on text conversion. Native speech-to-speech handling combined with asynchronous tool calling allows responsive conversational experiences without disruptive processing pauses. At the same time, ready-made integrations across providers like LiveKit and Agora lower the technical barrier for deploying production voice agents.

Perspectives

Coverage: 1× US · 1× Other

One story, several angles: how each source frames the topic, each with a verbatim quote.

Leaning: 1× Vendor PR

  • blog.googleVendor PRUS

    Google frames Gemini 3.8 Live as a major advancement in native speech-to-speech models that empowers developers to build real-time voice agents capable of multitasking during dialogue.

    Original quote

    Gemini 3.8 Live brings a step change to our native speech-to-speech models

    blog.google
  • buttondown.comOther

    The newsletter emphasizes that the model solves unnatural pauses in voice agents by reasoning and executing background tool calls in parallel with ongoing speech.

    Original quote

    Google's new Live API models hold a real-time voice conversation and keep reasoning in the background while they speak.

    buttondown.com

Source classification is maintained editorially (political spectrum only where consensus is broad; vendor communication is PR, not journalism). Unlabelled sources are unclassified: we do not guess.

Evidence

Solidly sourced
65/100
  • Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking in the Gemini Live API and Google AI Studio on September 15, 2026.

    verified
  • The new models feature native speech-to-speech processing with parallel background reasoning and asynchronous tool calling without conversational interruptions.

    single source
  • Infrastructure platforms including LiveKit, Agora, and Fishjam simultaneously announced integrations for Gemini 3.8 Live.

    verified
  • The technology targets entertainment use cases including interactive audio dramas, dynamic dubbing, and zero-latency AI radio co-hosts.

    single source

The evidence score is computed, not hand-set: from confidence, the number of sources and the share of verified statements.

Source & transparency

As of: September 17, 2026

AI-generatedAI-generated: produced automatically from vetted sources with technical quality checks (source, quote and figure verification); no human sign-off of each item before publication

Sources
2
Verified statements
2 / 4
Evidence score
65Solidly sourced

Want to put this into practice?

We connect you with suitable AI providers from the DACH region, free of charge and without obligation.

What's next?