Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism
Source
developer.nvidia.com
Author
Michelle Horton
Date
Why it matters
Turns transient GPU loss during long training runs into a recoverable event rather than a stall or restart, with resharding overhead under 1% — relevant to anyone whose training goodput is hostage to hardware flakiness.