
Artificial Analysis has measured 3,431 tokens/s on Gemma 4 31B via a private demonstration endpoint for NVIDIA Groq 3 LPX - the highest 100k-context output speed we have measured for Gemma 4 31B on either a private or public endpoint @nvidia announced today that the Groq 3 LPX ‘interactive AI inference’ rack is in full-scale production and will enter operation later this year. NVIDIA granted us early access to a private deployment of the rack serving Gemma 4 31B for benchmarking purposes We ran the standard 1k, 10k and 100k input sequence length prompts that we use for benchmarking serverless API providers. At both the 10k and 100k input sequence lengths, the endpoint delivered a median output speed of approx. 3,400 tokens/s, measured over 50 sequential (single-concurrency) requests. Output speed was maintained between the 10k and 100k input sequence lengths tested, indicating robust long-context inference performance Congratulations to the @nvidia team on the launch!
Long- output speed that holds steady from 10k to 100k input changes what interactive loops can afford, and the rack behind it enters service this year.
Checking sign-in…
Loading comments…