Shows how to allocate scarce GPU capacity without rewarding idle reservations: budgets, fair-share ordering and declared minimum runtimes. Ai2 reports 98% of owed GPU hours delivered at 98% occupancy.
Key takeaways · AI-distilled
Ai2's old priority scheduler bred two pathologies: researchers parked no-op jobs to squat GPUs because debug jobs could not launch fast enough, and priority inflation left 100% of workloads marked HIGH.
Every preemption-protected request must now be funded from budgets managers set down the program tree, so squatting spends the team's own allocation. The goal is making gaming costlier than arguing for a bigger budget.
Workloads declare a minimum runtime, capped at 8 hours, during which they cannot be preempted; afterward the scheduler may requeue them, which also lets unhealthy hosts drain automatically and cut human-in-the-loopRequiring a person's approval at specific points in an automated process, chosen so the irreversible steps are the ones a human sees.Full definition → repairs by 74%.
Debug jobs' p90 queue wait fell from 2 hours to 30 seconds, and on the largest H100 cluster median wait fell from 5 minutes to 24 seconds. 18% of delivered GPU time was unallocated, preemptible work filling gaps.
Interactive dev sessions got worse under the 8-hour protected cap, so Ai2 plans a CPU-only cluster and restorable sessions. It is also investigating fragmentation that may lengthen waits for the largest jobs.
Terms in this piece · Glossary
human-in-the-loop — Requiring a person's approval at specific points in an automated process, chosen so the irreversible steps are the ones a human sees.