Vibeleaderboard
← All Intel
Intel / post

Gemma 4 31B hits 3,400 tokens/s on NVIDIA's new inference rack

Source
Artificial Analysis
Date
Artificial Analysis@ArtificialAnlys

Artificial Analysis has measured 3,431 tokens/s on Gemma 4 31B via a private demonstration endpoint for NVIDIA Groq 3 LPX - the highest 100k-context output speed we have measured for Gemma 4 31B on either a private or public endpoint @nvidia announced today that the Groq 3 LPX ‘interactive AI inference’ rack is in full-scale production and will enter operation later this year. NVIDIA granted us early access to a private deployment of the rack serving Gemma 4 31B for benchmarking purposes We ran the standard 1k, 10k and 100k input sequence length prompts that we use for benchmarking serverless API providers. At both the 10k and 100k input sequence lengths, the endpoint delivered a median output speed of approx. 3,400 tokens/s, measured over 50 sequential (single-concurrency) requests. Output speed was maintained between the 10k and 100k input sequence lengths tested, indicating robust long-context inference performance Congratulations to the @nvidia team on the launch!

Terms in this piece · Glossary
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Long- output speed that holds steady from 10k to 100k input changes what interactive loops can afford, and the rack behind it enters service this year.

More from Artificial Analysis
Recommended reads
Comments

Checking sign-in…

Loading comments…