Olmo launches Olmo-core 3 — open, scalable MoE training infrastructure
Olmo has released Olmo-core 3, a redesigned open mixture-of-experts (MoE) training framework intended to scale MoE models into the trillion-parameter range while preserving computational efficiency. The stack replaces an FSDP-based implementation with a DDP-centered system, combines expert/pipeline parallelism and other optimizations, reports a 2.7× throughput uplift on a 47B-parameter MoE across eight NVIDIA B300 GPUs, and shows MXFP8 delivering about 21% higher throughput than BF16 in controlled tests.