Vibeleaderboard
← All Intel
Intel / post

OpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as its

Source
ArtificialAnlys
Date
From the Daily Brief

OpenAI's GPT-Live-1 debuted at the top of Artificial Analysis' Speech to Speech Index with a score of 81.5, edging out Grok Voice Think Fast 2.0 for the lead spot. The result comes from an architecture choice rather than a bigger voice model. GPT-Live-1 handles audio streaming and turn taking directly, then delegates reasoning and tool calls to a separately configured backend text model, splitting the job instead of asking one network to do both well. That separation lets OpenAI swap in a stronger text model for complex requests without retraining the voice component, and it explains why the model also cut interruption rates in earlier testing this month. For builders shipping real time voice agents, the pattern is now proven at the top of an independent benchmark. Keep the low latency audio loop thin, and hand anything that needs deep reasoning to a model built for that job.

Read the 2026-09-15 Brief →
ArtificialAnlys@ArtificialAnlys

OpenAI’s GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81.5 using Astra as its delegated backend model, ahead of Grok Voice Think Fast 2.0 GPT-Live-1 is @OpenAI's new full duplex Speech to Speech model that can delegate reasoning and tool use to a backend text model while continuing the conversation. Developers stream audio in and receive speech back through the API, with the backend text model configured separately. We evaluated two backend configurations: Astra at medium reasoning effort and Sol at low reasoning effort. Key takeaways: ➤ Speech to Speech Index: GPT-Live-1 (Astra, medium) achieves 81.5, ranking #1, while GPT-Live-1 (Sol, low) scores 80.1, ranking #3. Grok Voice Think Fast 2.0 High sits between them at 81.3 ➤ Speech Agent Arena: GPT-Live-1 (Sol, low) ranks #3 in preference at 1,053 Elo with 90.9% Task Success Rate, while GPT-Live-1 (Astra, medium) ranks #4 at 1,048 Elo with 87.4% task success. Gemini 3.1 Flash Live Minimal leads preference at 1,096 Elo, while Grok Voice Think Fast 2.0 High leads task success at 94.6% ➤ Tau Voice: GPT-Live-1 (Astra, medium) and GPT-Live-1 (Sol, low) take the top two spots on our…

Read the full post on X

Context

OpenAI's GPT-Live-1 is a full-duplex speech-to-speech model, meaning it can listen and speak at the same time rather than waiting for one side to finish. Instead of handling reasoning and tool calls itself, it delegates that work to a separately configured backend text model while the audio conversation continues, and Artificial Analysis tested two configurations: one backed by Astra at medium reasoning effort, one by Sol at low effort.

Artificial Analysis found the Astra-backed configuration debuted at #1 on its Speech to Speech Index at 81.5, ahead of xAI's Grok Voice Think Fast 2.0 High at 81.3, with the Sol-backed configuration close behind at 80.1. On the agentic-performance Tau Voice, both GPT-Live-1 configurations took the top two spots, ahead of Grok Voice. That top ranking did not last: Google DeepMind's Gemini 3.8 Live Extended Thinking, benchmarked by the same firm later the same day, scored 82.6 on the same index, edging past GPT-Live-1's debut score.

Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • tool use — A model's ability to call external functions — run code, search the web, edit files — instead of only generating text.
More from ArtificialAnlys
Recommended reads
Comments

Checking sign-in…

Loading comments…