A paper preview describes single-stage training with fixed hyperparameters and reports average accuracies of 54.9% and 64.3% on two 1.5B reasoning models.
Published on January 4, 2026, this entry summarizes JustRL, a single-stage reinforcement learning (RL) training approach with fixed hyperparameters, applied to two 1.5B reasoning models. The preview reports average accuracies of 54.9% and 64.3%; the entry does not specify tasks, evaluation sets, or experimental conditions.
The work offers a baseline for comparing simpler RL training recipes with more complex approaches. To verify the findings, consult the original paper and check its experimental setup, metrics, and evaluation sets before comparing accuracies. If you use AI to study or apply the material, do not submit confidential organizational data without authorization; use public or properly anonymized content.