A June 30, 2026 publication presents IW-OPD, a method that gives more weight to early tokens. The post reports faster convergence and a 6.9-point advantage over standard OPD on AIME-2025.
On June 30, 2026, researchers presented IW-OPD for on-policy distillation (OPD). The approach reweights tokens, assigning more weight to early ones and less to later ones, aiming to prevent the performance degradation attributed to later tokens in this process.
According to the post, the method converges faster and beats standard OPD by 6.9 points on AIME-2025. These are results reported by the publication; to assess the method, consult the original post and verify its experimental description, comparison conditions, and how the score was calculated.