Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
- ID
- 30788
- Status
- summarized
- Published
- 01 Oct 2026, 11:01 PM
- Fetched
- 01 Oct 2026, 11:23 PM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/allenai/olmocore3
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.0
- Created
- 01 Oct 2026, 11:24 PM
- Tags
- Audience
- developersai_ml_learners
What happened
Ai2 released Olmo-core 3, an open training framework for large mixture-of-experts models that replaces the earlier FSDP setup (gathering and resharding weights each batch) with DDP that keeps experts resident on GPUs and routes data to them. In one benchmark, growing the expert pool from 8 to 128 while still selecting 4 experts per token and holding active parameters near 3.2B raised total capacity from 4.6B to 47B with less than 5% throughput loss; the same stack was benchmarked past one trillion total parameters. A tech report, code, and interactive demo were published with it, and the post positions it against NVIDIA's Megatron-Core.
Why it matters
The usable number here is the ratio: roughly 10x total parameter capacity for under 5% throughput loss, which Ai2 attributes to the DDP resident-expert design rather than FSDP per-batch weight gathering. If you are picking a stack for any sparse/MoE training, that is the specific claim to reproduce on your own cluster before choosing Olmo-core 3 over Megatron-Core, because the routing and communication costs are what decide whether MoE actually saves you compute at your scale. For most readers who never train from scratch, the practical takeaway is narrower and honest: the open code and tech report document how expert-count scaling behaves, and the generation history (OlmoE at 64 routed experts, Olmo 3 dense, now this) shows Ai2 reversing its dense bet.
Discussion angle
Does the sub-5% throughput cost at 10x expert capacity survive outside Ai2's own setup, or does DDP resident-expert routing start losing to communication overhead once your interconnect and batch shapes differ?