Vibeleaderboard
← All Intel
Intel / article

Enhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor Parallelism

Source
developer.nvidia.com
Author
Michelle Horton
Date
Why it matters

Turns transient GPU loss during long training runs into a recoverable event rather than a stall or restart, with resharding overhead under 1% — relevant to anyone whose training goodput is hostage to hardware flakiness.

Read the source developer.nvidia.com
More from Michelle Horton
Recommended reads
Comments

Checking sign-in…

Loading comments…