Megatron-LM: model parallelism for large-model training
A 2019 paper describes splitting the weights of language models with billions of parameters across GPUs, addressing models whose weights do not fit on a single GPU.
The 2019 paper on Megatron-LM describes model parallelism for training language models with billions of parameters. The technique divides model weights across GPUs, addressing cases where the weights do not fit on a single GPU.
This Radar entry was published on April 23, 2026; that is the bookmark date, not the paper date. To verify the result, consult the original publication and check its date, scope, and account of the technique. If you use AI to study or apply the material, avoid submitting sensitive organizational data and follow your organization’s data-protection policies.