Efficient MoE training for biological foundation models using NVIDIA Transformer Engine and MXFP8
NVIDIA published a BioNeMo MoE training recipe that uses Transformer Engine primitives—GroupedLinear, MXFP8 block-scaled 8-bit precision, and a fused GroupedMLP kernel (ForwardGroupedMLP_CuTeGEMMSwiGLU_MXFP8)—to reduce kernel-launch overhead, lower activation memory, and speed MoE training. In an eight‑GPU (NVIDIA B200) benchmark the recipe reached up to 2.21× throughput versus a Hugging Face baseline; the fused MXFP8 kernel requires NVIDIA Blackwell hardware.