A comparison of six pruning methods for Llama-3.1-8B against training from scratch finds that the advantage depends on how the token budget is counted.
The study compares six pruning methods applied to Llama-3.1-8B with training models from scratch. The evaluations cover two token-budget scenarios.
When the comparison holds retraining tokens equal, pruned models have an advantage. When the total pipeline budget is counted instead, the result depends on pruning granularity.
To apply this kind of finding, first define which budget is being compared: retraining tokens alone or the whole pipeline. Then record the pruning granularity and use equivalent conditions when comparing alternatives.
If you use AI to study or apply the material, do not include personal data, secrets, or unnecessary internal content. Use synthetic or anonymized data where possible, and follow your organization’s access and retention rules.