NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt
Source
developer.nvidia.com
Author
Elizabeth Goodman
Date
Why it matters
AgentX measures inference the way agents actually load it — long prefill, KV-cache reuse, bursty concurrency — making it a better proxy for real serving cost than single-turn benchmarks.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
subagent — A separate agent spawned by another to do one scoped piece of work in its own context, returning only the result.
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.