A paper compares six pruning methods for Llama-3.1-8B with training from scratch under two token-budget scenarios. The advantage depends on how the budget is counted and on pruning granularity.
Published on July 5, 2026, the paper compares six pruning methods applied to Llama-3.1-8B with training models from scratch under two token-budget scenarios. The study examines alternatives for engineers assessing how to obtain smaller models.
With the same number of tokens for retraining, pruned models have an advantage. When the total pipeline budget is considered, however, the result depends on pruning granularity. Consult the original paper to check its methods, conditions, and full results; do not extend the conclusion to other models or budgets without verifying the data.