Vibeleaderboard
← All Intel
Intel / post

Gemini 3.8 Live Tops Speech-to-Speech Benchmarks

Source
ArtificialAnlys
Date
ArtificialAnlys@ArtificialAnlys

Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6% Gemini 3.8 Live is @GoogleDeepMind's successor to Gemini 3.1 Flash Live, a Speech to Speech model that executes tools and API calls in the background while continuing the conversation. It comes in two variants: the standard Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which supports configurable reasoning effort. We evaluated the standard model and the Extended Thinking variant at High reasoning effort through the Gemini Live API. Key takeaways: ➤ Speech to Speech Index: Gemini 3.8 Live Extended Thinking (High) debuts at #1 at 82.6, ahead of GPT-Live-1 (Astra, medium) at 81.5, Grok Voice Think Fast 2.0 High at 81.3 and GPT-Live-1 (Sol, low) at 80.1. The standard Gemini 3.8 Live debuts at #5 at 76.0, with both variants up on Gemini 3.1 Flash Live High at 71.5 (+11.1 and +4.5 points). The Index averages Speech Reasoning (Big Bench Audio), Agentic Performance (Tau Voice), Arena Preference and Arena Task Success Rate ➤ Speech Agent…

Read the full post on X

Context

Google DeepMind's Gemini 3.8 Live is a speech-to-speech model that executes tools and API calls in the background while a conversation continues, and it succeeds Gemini 3.1 Flash Live, according to Artificial Analysis's benchmark thread. It comes as a standard model and as Extended Thinking, which supports configurable reasoning effort.

Artificial Analysis tested Extended Thinking at High reasoning effort. It debuts first on the Speech to Speech Index at 82.6, ahead of GPT-Live-1 (Astra, medium) at 81.5 and Grok Voice Think Fast 2.0 High at 81.3, while the standard model is fifth at 76.0. It also leads the Tau Voice agentic at 68.6%, up from 37.7% for Gemini 3.1 Flash Live High. The lead has costs. On the Speech Arena, Extended Thinking trails on preference (Elo 990 versus 1083 for the standard model), its average time to first audio is 1.35 seconds versus 1.18, and it costs $3.50 per hour of input audio versus $0.84.

Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…