
Turns transient GPU loss during long training runs into a recoverable event rather than a stall or restart, with resharding overhead under 1% — relevant to anyone whose training goodput is hostage to hardware flakiness.
Checking sign-in…
Loading comments…