
Runway's platform team describes a working approach to balancing guaranteed capacity against idle GPUs on Kubernetes, using quota-based admission with cohort borrowing and preemption. They also identify the debugging trap that arises when admission and scheduling disagree.
Checking sign-in…
Loading comments…