Skip to content
Rota Nacional

Radar ·

Pre-training a 30B-A3B MoE model scaled from 16 to 512 B200 GPUs

Report of scaling with nearly linear scalability and 35.3% MFU, using small-scale tests to narrow the search space before adding more GPUs.

According to the Radar briefing published on 5 October 2026, an AI company described how it scaled the pre-training of a 30B-A3B MoE model from 16 to 512 B200 GPUs. The report states nearly linear scalability and an MFU (model FLOPs utilization) of 35.3%.

The central approach is to reduce the search space at small scale, testing training configurations before increasing the number of GPUs. The text presents this as a practical way to validate choices before allocating more hardware. The full methodology is in the original, which should be consulted to confirm the figures and the test conditions.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 10,00.

Try free