Vibeleaderboard
← All Intel
Intel / article

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Source
allenai.org
Date
Key takeaways · AI-distilled
  • Ai2 grew an 's expert pool from 8 to 128 while still routing each to four experts, holding active parameters near 3.2B. Total parameters rose from 4.6B to 47B and training throughput fell by less than 5%.
  • Olmo-core 3 drops the FSDP setup that gathered and resharded weights for every small batch in favor of a DDP-based design that keeps experts resident on GPUs. In a preliminary test on eight B300s, a 47B MoE hit 52,000 tokens/s per GPU versus 19,400 before, about 2.7x.
  • Enabling MXFP8 where it helped most raised throughput about 21% over BF16 on four B300s and cut peak active memory from 103 GiB to 95 GiB. Ai2 says most of the gain came from feed-forward compute and moving data between experts, not .
  • A 1.2-trillion-parameter configuration with 58.36B active parameters on 512 GPUs peaked at 858 TFLOP/s per GPU. Those runs used random routing, so they measure system performance rather than the quality of a trained model.
  • Ai2's report names a failure it calls token gerrymandering: a balanced-routing score improved while the real workload grew less balanced. It also found that overlapping communication with compute on separate GPU streams sometimes slowed training.
Terms in this piece · Glossary
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
Why it matters

Olmo-core 3 is open training infrastructure for large MoE models, with code and a tech report. Teams studying or training MoEs can inspect how expert routing and communication costs are handled at scale.

Read the source allenai.org
Recommended reads
Comments

Checking sign-in…

Loading comments…