What It Means to 'Converse' with AI: The Landscape Shifted by Gemini 3.8 Live

An illustration depicting a person and an AI voice assistant having a natural conversation via a smartphone
AI Summary

Google's latest voice models, Gemini 3.8 Live and Extended Thinking, significantly enhance real-time voice interaction, allowing AI to perform complex tasks without interrupting the flow of conversation.

Imagine this: You wake up in the morning and say to your smartphone, “I have a lot of meetings today; find the closest restaurant for lunch and book a table for me.” With previous AI assistants, you would have faced long pauses as the AI searched and booked—or worse, the conversation would simply cut off. But now, we are entering a world where AI acts like a capable assistant sitting right next to you, continuing the conversation seamlessly: “I’m looking up restaurants now. I found a few good ones; shall I go with one where you can book for 12:30?”

Google recently unveiled ‘Gemini 3.8 Live’ and ‘Gemini 3.8 Live Extended Thinking,’ the most advanced voice conversation models yet to make this future a reality Source 1 Source 4.

Why Is This Important?

Until now, many voice AIs we used created a sense of “dead air” while searching data and generating answers. It was like a chef who went completely silent in the kitchen while cooking, only reappearing once the dish was finished.

However, these newly released models were designed specifically for ‘real-time voice agents’ Source 3. They allow the AI to perform complex tasks while speaking with the user, ensuring the context of the conversation is never lost Source 1. This means a much more natural and deeper level of collaboration with AI is now possible.

Understanding the Technology

To explain this tech, let’s use two analogies.

First, ‘Talking While Thinking’: The ‘Gemini 3.8 Live Extended Thinking’ model is similar to how we talk to ourselves while pondering a problem. When the AI processes a complex request, it provides progress updates via voice, such as “I’m looking through the data right now” or “This information is taking a little time to process” Source 8. This lets the user know the AI hasn’t given up, but is working hard.

Second, ‘A Wizard with Eyes and Ears’: These models don’t just listen to your voice. They are ‘multimodal’ AIs that can simultaneously understand voice, images, video, and text Source 2. It’s like having a person who sees and hears everything participate in a conversation to make comprehensive judgments. This is possible because they can remember and analyze a vast amount of information—up to 128K tokens—at once Source 2.

Current Status

Google’s models are already delivering impressive results. According to evaluations by the organization ‘Artificial Analysis,’ the ‘Gemini 3.8 Live Extended Thinking’ model took first place overall in the ‘Speech-to-Speech Quality Index’ Source 10. Recording a high score of 82.6, it has proven itself to be at the industry’s cutting edge Source 7 Source 12.

Furthermore, because they support over 97 languages globally, users—not just Korean speakers, but people around the world—can communicate more conveniently with their AI assistants Source 11 Source 15. You can experience this technology right now via the Gemini app or Google AI Studio Source 8.

What Lies Ahead?

Over the past few months, Google has been updating its core Gemini models almost every three weeks Source 5. This demonstrates how quickly AI is integrating into the everyday devices we use.

In the future, AI assistants will evolve beyond simple command-receivers into true ‘partners’ who think and advise alongside us as we work or sort through dilemmas. This update marks a significant milestone where conversations with AI move beyond mere technical communication and closer toward human-like interaction.

MindTickleBytes AI Reporter’s View

Human conversation is not merely an exchange of information, but a process of sharing a ‘flow of thought.’ This model brings us one step closer to a level where AI can ponder along with humans without disrupting this flow.

References

  1. Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking - The Keyword
  2. Gemini 3.8 Audio (Live, Live Extended Thinking) - Model Card
  3. Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production-Grade Voice Agents
  4. Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail
  5. Gemini 3.8 Live & Extended Thinking: Google Voice AI [2026]
  6. Gemini 3.8 Live Launch: 82.6 Voice AI Score [2026]
  7. Meet Gemini 3.8 Live and 3.8 Live Extended Thinking: our best…
  8. 「Gemini 3.8 Live/3.8 Live Extended Thinking」正式発表 – Jetstream
  9. [Gemini 3.8 Live и 3.8 Live Extended Thinking… AiManual](https://ai-manual.ru/article/gemini-38-live-i-38-live-extended-thinking-golosovyie-ai-agentyi-s-parallelnyim-myishleniem/)
  10. Google DeepMind выпустила Gemini 3.8 Live и Extended Thinking…
  11. Gemini 3.8 Live: Google Splits Its Voice Line in Two
  12. Gemini 3.8 Live Extended Thinking Pricing, Specs & Sources
  13. Google launches Gemini 3.8 Live to take on OpenAI’s GPT-Live-1 at…
AD
Test Your Understanding
Q1. What is the most significant feature of the newly announced 'Extended Thinking' model?
  • It can only recognize images
  • It maintains the flow of conversation by explaining its own progress
  • It can only communicate via text
The Extended Thinking model features enhanced capabilities that allow the AI to speak about its thoughts or task progress while performing complex work, ensuring the conversation remains uninterrupted.
Q2. What types of input data can Gemini 3.8 Live models process?
  • Voice data only
  • Text only
  • Voice, images, video, and text
The new models support multimodal input, including voice, images, video, and text.
Q3. What overall score did this model achieve in Artificial Analysis's Speech-to-Speech rankings?
  • 70.5
  • 82.6
  • 90.0
The Gemini 3.8 Live Extended Thinking model achieved an overall first-place ranking with a score of 82.6 on Artificial Analysis's voice conversation quality index.
What It Means to 'Converse'...
0:00