Will conversations with AI become more natural? Google's new voice AI models, the story of Gemini 3.8 Live

A visual graphic representing Google's new real-time conversational AI model, Gemini 3.8 Live.
AI Summary

Google's newly announced real-time voice-based conversational AI, 'Gemini 3.8 Live,' and 'Extended Thinking' models are designed to make AI conversations feel like smooth, intelligent, real-world dialogues.

Imagine this: You wake up on a busy morning and ask your smartphone’s AI assistant, “Can you organize the materials for this afternoon’s meeting and summarize them for me as if we were talking?” and the conversation continues naturally, as if a very smart assistant were right beside you answering. Moving beyond the rigid command input methods of the past, Google has taken a major step toward a seamless experience that feels like talking to a human.

On September 15, 2026, Google officially announced its new voice-centric AI models, ‘Gemini 3.8 Live’ and ‘Gemini 3.8 Live Extended Thinking’ [Source: Google Launches Gemini 3.8 Live and Extended Thinking Models]. AI is now evolving from a machine that simply executes commands into our daily conversational partner.

Why is this important?

The voice AI we have experienced so far felt like an automated response system (ARS) reading a menu aloud. It would often stutter if you interrupted it or asked a complex question. However, the models announced this time aim to provide a smooth experience that does not disrupt the flow of conversation with the user, making it feel as if you are communicating with a real person [Source: Google Releases Gemini 3.8 Live-Extended Conversational Model].

For developers and companies, this means they can more easily implement much more familiar and efficient voice services for users, such as smart customer support bots or personal assistant services [Source: Google Launches Gemini 3.8 Live and Extended Thinking Voice Models].

Understanding it easily

To help understand this new technology, let’s use two metaphors.

Simply put, the first is the ‘continuous film’ metaphor. While existing AI models processed photos one by one separately, Gemini 3.8 Live models process audio, video, and text data in real-time without interruption, like a long movie film [Source: Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card]. This allows them to understand the situation immediately while we are speaking and react naturally. It’s like actors exchanging lines seamlessly in a movie.

The second is the ‘depth of thought’ metaphor. This is why the models were divided into two. If ‘Gemini 3.8 Live’ is a ‘fast and agile athlete’ optimized for daily, efficient conversation, you can easily think of ‘Gemini 3.8 Live Extended Thinking’ as a ‘thoughtful analyst’ that handles conversations requiring complex problems or deep logical thinking [Source: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking]. It’s similar to an expert who pauses for a moment to think deeply before answering when solving difficult math problems or creating complex plans.

Where can it be used?

Currently, these technologies have been made public so that developers and companies can utilize them directly through the Gemini API (Application Programming Interface), Google AI Studio, and Gemini Enterprise [Source: Google Launches Gemini 3.8 Live and Extended Thinking Voice Models]. This means it is highly likely that we will soon encounter smarter voice AI in the app services we use in our daily lives.

What’s next?

In the future, when we talk to AI, the feeling of ‘collaboration’—asking for and receiving necessary information—will likely be stronger than the feeling of simply giving ‘orders’. Since they can recognize information combining video and audio in real-time, we expect to see more advanced agent services, such as an AI explaining or thinking along with the situations shown through a smartphone camera in real-time [Source: Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card].

We seem to have moved past the stage where we treated AI as a tool for inputting commands, bringing us a little closer to an era of partnership where we share our daily lives with AI.

AI’s Take

Google DeepMind describes these models as “the best conversational AI that talks, thinks, and processes tasks in the background without breaking the user’s flow of conversation” [Source: Google DeepMind on X].

The core of the Gemini 3.8 Live model is that it has secured the ‘context of conversation’ beyond the ‘speed of technology’. Rather than how fast an AI can answer, how naturally it can blend in with us is becoming the true benchmark of innovation.

References

  1. Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking
  2. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card
  3. Gemini 3.8 Live: gemini-3.8-live, 97 языков и Extended Thinking
  4. Google Launches Gemini 3.8 Live and Extended Thinking Models
  5. Google Releases Gemini 3.8 Live-Extended Conversational Model
  6. Google DeepMind on X
  7. Google Launches Gemini 3.8 Live and Extended Thinking Voice Models
  8. Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
AD
Test Your Understanding
Q1. What is the primary purpose of the Gemini 3.8 Live models announced by Google?
  • Improve image generation speed
  • Implement more natural and intelligent real-time voice conversations
  • Build a dedicated text translation engine
The purpose of these models is to make conversations with AI more flexible and intelligent, making them feel like talking to an actual person.
Q2. What kind of data do Gemini 3.8 Live models process?
  • Text data only
  • Audio data only
  • Real-time processing of continuous audio, video, and text streams
These models process audio, video, and text as a continuous stream in real-time, providing immediate voice responses.
Q3. Where can the Gemini 3.8 Live models be used?
  • Available via Gemini API, Google AI Studio, etc.
  • Only works on dedicated hardware devices
  • Operates only in offline environments
These models are provided to developers and enterprises through Google's Gemini API, Google AI Studio, and Gemini Enterprise.
Will conversations with AI ...
0:00