inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
Shows how to get bin-packing and controlled spill for GPU workloads without forking the Kubernetes scheduler, using taints driven by a PromQL signal plus the Descheduler and Kueue for admission.