HomeNews

Gemini 3.8 Live and Extended Thinking Launch

Madhav Jha
Madhav Jha
Sep 22, 2026 1:54 AM
0
 min read
Select Emergent as your Preferred news source
Gemini 3.8 Live and Extended Thinking Launch

💡 TL;DR

  • Google DeepMind launched Gemini 3.8 Live and 3.8 Live Extended Thinking on September 15, 2026, bringing real-time voice conversations to the Gemini ecosystem.
  • The Extended Thinking variant adds advanced reasoning capabilities, allowing the model to pause and work through multi-step problems before responding to complex queries.
  • Both models integrate seamlessly with Google's ecosystem, targeting education, accessibility, and professional workflows requiring conversational AI with deep reasoning.

Google DeepMind has introduced two new conversational AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Officially launched on September 15, 2026, these models bring real-time voice interaction and advanced reasoning to the Gemini family, extending the capabilities first seen in earlier Flash iterations. The release targets users who need hands-free, intelligent assistance for tasks ranging from education to professional workflows.

What Makes Gemini 3.8 Live Different

Gemini 3.8 Live is optimized for natural, real-time voice conversations. Unlike text-based models, it processes spoken input and responds instantly, making it suitable for scenarios where typing is impractical: driving, cooking, or accessibility use cases. The model leverages the efficiency of the Gemini 3.8 Flash architecture while adding low-latency audio processing. Early tests show response times under 500 milliseconds for most queries, a significant improvement over previous voice assistants.

The model integrates with Google's ecosystem, including Workspace, Assistant, and Android devices. Developers can access it via the Gemini API, with pricing structured around audio tokens rather than text tokens. This shift reflects the model's focus on spoken interaction, where brevity and clarity matter more than raw token count.

Extended Thinking Adds Reasoning Depth

Gemini 3.8 Live Extended Thinking introduces a pause-and-reason capability. When faced with complex multi-step problems, the model explicitly takes time to work through intermediate steps before delivering an answer. This is similar to OpenAI's o1 series but adapted for voice interactions. For example, when asked to plan a multi-city trip with budget constraints, the model might say, "Let me think about this," then outline its reasoning process aloud before presenting the final itinerary.

This extended reasoning mode is optional. Users can toggle it on for tasks requiring careful analysis, such as coding assistance, financial planning, or scientific problem-solving. In standard mode, the model prioritizes speed. In extended mode, it sacrifices speed for accuracy, often taking 5 to 15 seconds to formulate responses. Internal benchmarks show a 30 percent improvement in multi-step reasoning tasks compared to the base Gemini 3.8 Live model.

Target Use Cases and Availability

Google positions Gemini 3.8 Live for three core audiences: students seeking tutoring, professionals needing hands-free productivity tools, and users with accessibility needs. The Extended Thinking variant is particularly relevant for STEM education, where walking through problem-solving steps aloud can reinforce learning. Early adopters include universities using it for interactive study sessions and enterprise teams integrating it into voice-enabled dashboards.

Both models are available now through Google AI Studio and the Gemini API. Pricing starts at $0.05 per 1,000 audio tokens for standard Live mode, with Extended Thinking priced at $0.12 per 1,000 tokens to account for the additional compute overhead. Free-tier users get 50 queries per day in standard mode, 10 in extended mode.

Comparison to Existing Gemini Models

Gemini 3.8 Live builds on the Flash family's efficiency gains. Compared to Gemini 3.7 Flash, it delivers 40 percent faster audio processing while maintaining competitive performance on text benchmarks. The Extended Thinking mode represents a new direction for Google, acknowledging that speed is not always the priority. While Gemini 3.8 Flash benchmarks emphasize throughput, these Live variants emphasize user experience in conversational settings.

  • Gemini 3.8 Live: Optimized for real-time voice, sub-500ms latency, no extended reasoning.
  • Gemini 3.8 Live Extended Thinking: Adds pause-and-reason capability, 5-15 second response times for complex queries.
  • Both models: API-accessible, priced by audio token, integrated with Google ecosystem.

What This Means

The launch of Gemini 3.8 Live and Extended Thinking signals Google's commitment to multimodal AI beyond text. By prioritizing voice and reasoning, the company addresses gaps left by earlier assistants that lacked depth or speed. For developers, these models offer new primitives for building conversational apps. For end users, they represent a step toward AI that listens, thinks, and responds in ways that feel natural. As competition intensifies among frontier labs, Google's focus on practical, accessible interfaces may prove as important as raw benchmark scores.

About the writer

Madhav is the Co-founder and CTO at Emergent, leading the technical vision behind the platform's autonomous coding agents that power production-ready app development for millions of users worldwide.

Start Building
on Emergent today
Try Emergent