The DAPO paper reports that standard GRPO scored 30% on AIME in its experiments; adding some techniques raised the reported result to 50%.
Published on March 1, 2026, the report on the DAPO paper compares standard GRPO with a configuration using some additional techniques on AIME: 30% and 50%, respectively. These figures are results reported by the authors, not a guarantee of performance in other settings.
The comparison may interest engineers evaluating reinforcement-learning training methods. To verify the finding, consult the original paper and check its experimental setup, added techniques, and evaluation conditions before comparing or reproducing the results.