Summaries
Short AI and tech summaries with source links, signal scores, and why each update matters for builders, founders, and Malaysian tech workers.
Showing 1-1 of 1 results
| Date | Provider | Score | Summary |
|---|---|---|---|
| 01 Oct 2026, 11:01 PM | Hugging Face Blog | 6.0 | Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Ai2 released Olmo-core 3, an open training framework for large mixture-of-experts models that replaces the earlier FSDP setup (gathering and resharding weights each batch) with DDP that keeps experts resident on GPUs and routes data to them. In one benchmark, growing the expert pool from 8 to 128 while still selecting 4 experts per token and holding active parameters near 3.2B raised total capacity from 4.6B to 47B with less than 5% throughput loss; the same stack was benchmarked past one trillion total parameters. A tech report, code, and interactive demo were published with it, and the post positions it against NVIDIA's Megatron-Core. Why: The usable number here is the ratio: roughly 10x total parameter capacity for under 5% throughput loss, which Ai2 attributes to the DDP resident-expert design rather than FSDP per-batch weight gathering. If you are picking a stack for any sparse/MoE training, that is the specific claim to reproduce on your own cluster before choosing Olmo-core 3 over Megatron-Core, because the routing and communication costs are what decide whether MoE actually saves you compute at your scale. For most readers who never train from scratch, the practical takeaway is narrower and honest: the open code and tech report document how expert-count scaling behaves, and the generation history (OlmoE at 64 routed experts, Olmo 3 dense, now this) shows Ai2 reversing its dense bet. |