Skip to content
Rota Nacional

Radar ·

RLoop alternates RL and policy consolidation

Published on November 7, 2025, the account describes an iterative training strategy for RLVR and reported gains in a math evaluation.

Published on November 7, 2025, the account presents RLoop, which alternates reinforcement-learning (RL) exploration with consolidation through RFT from filtered expert trajectories. The proposal is a policy initialization strategy for RLVR pipelines, intended to reduce overfitting and forgetting.

In the described experiment, using Qwen-2.5-7B-Math with DAPO-17K, the post reports about 9% higher Avg@32 and more than 15% higher Pass@32 compared with RL. These are results reported by the source, not a guarantee of general performance. To assess the claim, consult the original post and check its setup, metrics, and comparison conditions; the available summary does not detail them.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free