A study trains a hint generator and a hint-conditioned solver with reinforcement learning. The report says it achieved 44% higher math accuracy on AIME 2025 than RL with long chains.
The paper presents a hint generator and a solver that uses those hints, both trained with reinforcement learning. It tests whether short, reusable guidance can support reasoning tasks without relying on longer chains of thought.
According to the report, the method achieved 44% higher math accuracy on AIME 2025 than RL with long chains. This result applies to the stated comparison and does not guarantee gains on other tasks. When studying or applying the work with AI, avoid submitting personal data, confidential code, or internal documents; use synthetic or anonymized examples.