Vibeleaderboard
Index / article

How to Train Really Large Models on Many GPUs?

lilianweng.github.io
Visit lilianweng.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

[Updated on 2022-03-13: add expert choice routing .] [Updated on 2022-06-10]: Greg and I wrote a shorted and upgraded version of this post, published on OpenAI Blog: “Techniques for Training Large Neural Networks”

Why it made the leaderboard

A rigorous, well-organized reference on the parallelism strategies (data, tensor, pipeline, MoE, expert-choice routing) and memory-saving techniques that make large-scale model training feasible — useful for anyone reasoning about distributed training tradeoffs.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.