Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RELEASE · MODELS · #1359

Olmo launches Olmo-core 3 — open, scalable MoE training infrastructure

Olmo has released Olmo-core 3, a redesigned open mixture-of-experts (MoE) training framework intended to scale MoE models into the trillion-parameter range while preserving computational efficiency. The stack replaces an FSDP-based implementation with a DDP-centered system, combines expert/pipeline parallelism and other optimizations, reports a 2.7× throughput uplift on a 47B-parameter MoE across eight NVIDIA B300 GPUs, and shows MXFP8 delivering about 21% higher throughput than BF16 in controlled tests.

KEY POINTS

  1. Olmo has released Olmo-core 3, a redesigned open mixture-of-experts (MoE) training framework intended to scale MoE models into the trillion-parameter range while preserving computational efficiency.
  2. The stack replaces an FSDP-based implementation with a DDP-centered system, combines expert/pipeline parallelism and other optimizations, reports a 2.7× throughput uplift on a 47B-parameter MoE across eight NVIDIA B300 GPUs, and shows MXFP8 delivering about 21% higher throughput than BF16 in controlled tests.
  3. An open, benchmarked MoE training stack that meaningfully improves throughput and memory efficiency can lower the cost and technical barrier to training very large sparse models.

WHY IT MATTERS

An open, benchmarked MoE training stack that meaningfully improves throughput and memory efficiency can lower the cost and technical barrier to training very large sparse models.

SOURCES & TIMELINE

1