Vibeleaderboard
← All Intel
Intel / post

Gemini 3.5 Transcribe measured: 2.6% WER, and a live variant at 0.40s

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys
Thread · 5 parts

Google has released Gemini 3.5 Transcribe, ranking #5 on AA-WER at 2.6%, alongside Gemini 3.5 Transcribe Live, achieving 4.0% AA-WER Streaming at 0.40s after speech end @GoogleAI's Gemini 3.5 Transcribe release comprises two API offerings: Gemini 3.5 Transcribe Live for continuous streaming through the Live API, and Gemini 3.5 Transcribe for pre-recorded audio through the Interactions API. Both support 85+ languages, custom vocabulary and automatic text formatting, while the pre-recorded API adds timestamped multi-speaker identification for up to three speakers, with support beyond three currently experimental. Key takeaways ➤ Non-streaming transcription: Gemini 3.5 Transcribe achieves 2.6% AA-WER, ranking #5 overall, and processes audio at approximately 84× realtime. ➤ First Final Transcription: Gemini 3.5 Transcribe Live achieves 4.0% WER, with its first final-denoted transcript arriving 0.40s after VAD-detected end of speech. ➤ First Partial Transcription: Gemini 3.5 Transcribe Live achieves 5.8% WER, with its first transcript-bearing event arriving 0.25s after detected end of speech. ➤ Price: Gemini 3.5 Transcribe costs approximately $5 per 1,000 minutes and Transcribe Live $9, assuming 25 audio tokens per second and 175 text tokens per minute at their respective API rates ($2 per 1M audio input tokens and $12 per 1M text output tokens for Transcribe; $3.50 and $21, respectively, for Live). See more details below ⬇️

Gemini 3.5 Transcribe processes approximately 84 seconds of audio per second and costs approximately $5 per 1,000 minutes. Among other models in the top five for non-streaming accuracy, it is faster than ElevenLabs Scribe v2 at 55× realtime, although slower than Microsoft MAI-Transcribe-1.5 at 191× and Smallest AI Pulse Pro at 275×. It is less expensive than MAI-Transcribe-1.5 at $6 per 1,000 minutes, but more expensive than Scribe v2 at $3.67 and Smallest AI Pulse Pro at $4.

On First Partial Transcription, Gemini 3.5 Transcribe Live achieves 5.8% WER at 0.25s after detected end of speech. GPT Live Transcribe achieves 6.3% WER at 0.26s, making Gemini slightly faster and more accurate at this measurement point.

Gemini 3.5 Transcribe Live costs approximately $9 per 1,000 minutes, based on Google’s estimate of 25 audio input tokens per second and 175 text output tokens per minute, with API pricing of $3.50 per 1M input tokens and $21 per 1M output tokens. This is higher than Cartesia Ink-2 at $4, Deepgram Nova-3 Realtime at $4.80 and ElevenLabs Scribe v2 Realtime at $6.50.

Read the full thread on X
Terms in this piece · Glossary
  • streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Accuracy, latency and price figures for Gemini 3.5 Transcribe and its Live variant, compared against Scribe v2, Deepgram Nova-3 and MAI-Transcribe, before you commit a voice to a speech backend.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…