Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Source
allenai.org
Date
Key takeaways · AI-distilled
Ai2 grew an mixture-of-expertsA model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.Full definition →'s expert pool from 8 to 128 while still routing each tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition → to four experts, holding active parameters near 3.2B. Total parameters rose from 4.6B to 47B and training throughput fell by less than 5%.
Olmo-core 3 drops the FSDP setup that gathered and resharded weights for every small batch in favor of a DDP-based design that keeps experts resident on GPUs. In a preliminary test on eight B300s, a 47B MoE hit 52,000 tokens/s per GPU versus 19,400 before, about 2.7x.
Enabling MXFP8 where it helped most raised throughput about 21% over BF16 on four B300s and cut peak active memory from 103 GiB to 95 GiB. Ai2 says most of the gain came from feed-forward compute and moving data between experts, not attentionThe mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.Full definition →.
A 1.2-trillion-parameter configuration with 58.36B active parameters on 512 GPUs peaked at 858 TFLOP/s per GPU. Those runs used random routing, so they measure system performance rather than the quality of a trained model.
Ai2's report names a failure it calls token gerrymandering: a balanced-routing score improved while the real workload grew less balanced. It also found that overlapping communication with compute on separate GPU streams sometimes slowed training.
Terms in this piece · Glossary
mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
attention — The mechanism that lets a model weigh which earlier words matter for the word it's currently processing — the core operation of a transformer.
Why it matters
Olmo-core 3 is open training infrastructure for large MoE models, with code and a tech report. Teams studying or training MoEs can inspect how expert routing and communication costs are handled at scale.