Vibeleaderboard
← All Intel
Intel / article

How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency

Source
Sarah McKenney
Author
Sarah McKenney
Date
Key takeaways · AI-distilled
  • Static planning reserves peak power for every node at once, but training cycles through compute, communication and checkpointing and through prefill, decode and idle gaps, so headroom inside one node's reservation stays stranded while the site sits below its limit.
  • NVIDIA stresses MaxLPS adds no site power: operator policy sets node and group limits, reserves and priorities, and the software shifts participating GPU power caps from telemetry so the aggregate stays under the approved budget.
  • In the NVIDIA and Nscale test, the extra 52 GPUs ran one more high-throughput Kimi K2.5 instance while per-instance output held flat (59,153 vs 59,220 /s); measured power rose 19.7% and budget utilization went from 62.9% to 75.2%.
  • Usable headroom depends on the workload mix: workloads with complementary power profiles free more capacity than ones that peak together, so NVIDIA advises testing a representative production mix against the aggregate limit rather than a single model.
  • The recommended rollout is staged: map the enforceable power boundary, baseline under static provisioning, start with limits near that baseline, add nodes in steps while testing peak demand and telemetry failure, then set production limits.
Terms in this piece · Glossary
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters

NVIDIA and Nscale show that policy-governed power sharing (DSX MaxLPS) let them run 192 GPUs instead of 140 within an identical power budget, lifting throughput per watt from 4.10 to 6.12 tokens/s/W, though P99 time-to-first-token rose 17%.

Read the source developer.nvidia.com
Recommended reads
Comments

Checking sign-in…

Loading comments…