Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar
Source
Bryan Shan
Author
Bryan Shan
Date
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
SemiAnalysis's own AgentX benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → shows Rubin NVL72 delivering up to 7x better agentic tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → throughput per megawatt than Blackwell on early software, implying roughly 2x profit per gigawatt as the stack matures.