
Distributed-training bugs are silent gradient bugs; this shows the concrete failure of naive chunk/all_gather sharding and what DTensor's placement propagation does and does not buy you at scale.
articleTraining Variable Long Sequences with Data-Centric ParallelGeng Zhang, Xuanlei Zhao, Kai Wang, Yang YouChecking sign-in…
Loading comments…