Pre-training a 30B-A3B MoE model scaled from 16 to 512 B200 GPUs
Report of scaling with nearly linear scalability and 35.3% MFU, using small-scale tests to narrow the search space before adding more GPUs.
According to the Radar briefing published on 5 October 2026, an AI company described how it scaled the pre-training of a 30B-A3B MoE model from 16 to 512 B200 GPUs. The report states nearly linear scalability and an MFU (model FLOPs utilization) of 35.3%.
The central approach is to reduce the search space at small scale, testing training configurations before increasing the number of GPUs. The text presents this as a practical way to validate choices before allocating more hardware. The full methodology is in the original, which should be consulted to confirm the figures and the test conditions.