A practical discussion of FP4 training for large-scale Mixture-of-Experts models on Hopper GPUs identifies activation memory and expert-parallel communication as bottlenecks.
The material covers practical FP4 training of large-scale Mixture-of-Experts (MoE) models on Hopper GPUs. It identifies two bottlenecks: memory used by activations and communication among experts in the expert-parallel scheme.
These issues matter to engineers studying memory use and communication in large-scale MoE training. If you use AI to study or apply the material, do not submit internal data, credentials, or personal information; use fictional or anonymized examples instead.