Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

NEWS · CODING · #964

Efficient MoE training for biological foundation models using NVIDIA Transformer Engine and MXFP8

NVIDIA published a BioNeMo MoE training recipe that uses Transformer Engine primitives—GroupedLinear, MXFP8 block-scaled 8-bit precision, and a fused GroupedMLP kernel (ForwardGroupedMLP_CuTeGEMMSwiGLU_MXFP8)—to reduce kernel-launch overhead, lower activation memory, and speed MoE training. In an eight‑GPU (NVIDIA B200) benchmark the recipe reached up to 2.21× throughput versus a Hugging Face baseline; the fused MXFP8 kernel requires NVIDIA Blackwell hardware.

KEY POINTS

  1. NVIDIA published a BioNeMo MoE training recipe that uses Transformer Engine primitives—GroupedLinear, MXFP8 block-scaled 8-bit precision, and a fused GroupedMLP kernel (ForwardGroupedMLP_CuTeGEMMSwiGLU_MXFP8)—to reduce kernel-launch overhead, lower activation memory, and speed MoE training.
  2. In an eight‑GPU (NVIDIA B200) benchmark the recipe reached up to 2.21× throughput versus a Hugging Face baseline; the fused MXFP8 kernel requires NVIDIA Blackwell hardware.
  3. Fused grouped-expert kernels and block-scaled 8-bit training can materially improve GPU utilization and memory efficiency, enabling larger and longer-sequence MoE biological models to be trained more cost‑effectively on supported NVIDIA hardware.

WHY IT MATTERS

Fused grouped-expert kernels and block-scaled 8-bit training can materially improve GPU utilization and memory efficiency, enabling larger and longer-sequence MoE biological models to be trained more cost‑effectively on supported NVIDIA hardware.

SOURCES & TIMELINE

1