Google makes Gemini 3.8 Live generally available for voice agents

Google's official Gemini 3.8 Live launch image.Google
Google's official Gemini 3.8 Live launch image.Google
AI & Automation

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking for real-time voice applications, with access through the Gemini API and Vercel AI Gateway.

Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking for real-time voice applications. The models combine spoken conversation with visual grounding, tool use, and—on the Extended Thinking variant—parallel multi-step reasoning. Google says both are available now through the Gemini API, while Vercel has added them to its AI Gateway for applications built with the AI SDK.

Gemini 3.8 Live targets low-latency voice workflows

The standard Gemini 3.8 Live model is designed for scalable, fluid conversations. Google's announcement highlights real-time visual context, language support, and background tool calls that continue while the user keeps speaking. Vercel's implementation identifies the model as google/gemini-3.8-live and documents automatic switching across 97 languages.

That combination fits assistants that need to listen, inspect context, and call services without forcing a turn-based chat pattern. Examples include spoken customer support, guided interfaces, accessibility features, and applications that can acknowledge a request while work continues in the background.

Extended Thinking adds reasoning during the conversation

Gemini 3.8 Live Extended Thinking is aimed at higher-complexity tasks. Google describes multi-step reasoning that runs in parallel with speech; Vercel similarly says the model can acknowledge a request and narrate progress without interrupting the interaction. That is a meaningful distinction for voice agents handling research, planning, or tool-driven workflows where silence can feel like failure.

Concept illustration: Extended Thinking adds reasoning during the conversation
AI-generated illustration

Google also reports benchmark results for its Extended Thinking model, including a 82.6 score on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice agentic task completion, 35.1% on Sierra's τ-Voice-banking benchmark, and 97.7% on Big Bench Audio. These are vendor-reported results from the launch material, not an independent LinkLoot test, and the announcement notes that the EVA-Bench run used the Live API on Gemini Enterprise Agent Platform.

Access through Gemini API and Vercel AI Gateway

Developers can use the models through Google's Gemini API and Google AI Studio. Vercel's AI Gateway exposes both models through its realtime API, with short-lived tokens and WebSocket-based event handling. The Gateway route adds a vendor-neutral integration surface with usage and cost tracking, retries, failover, and performance controls, according to Vercel's changelog.

Teams evaluating the release should separate model capability from integration overhead. Realtime voice systems still need interruption handling, transcript persistence, tool authorization, latency monitoring, and explicit treatment of generated-audio disclosure. Google says the models are generally available, but production teams should confirm current quotas, pricing, regional availability, and the exact terms for their chosen endpoint before committing to a rollout.

For implementation patterns around agent orchestration, see LinkLoot's AI workflow automation guide.

Evidence and rollout status

Google's launch page is the primary announcement and states that Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available today. Vercel independently confirms the model IDs and their availability on AI Gateway. Together, the sources establish a current API and integration availability milestone; they do not establish identical quotas or pricing across Google's direct API, AI Studio, and Vercel's gateway.

From reading to doing

Try the related loot

Give OpenClaw Agents Pre-Verified Web Actions with Actionbook

Open loot