On-policy distillation routes supervision per token
A paper describes a method that dynamically selects, token by token, between supervision based on a teacher and on the model itself.
The paper presents Dual On-Policy Distillation, an approach that dynamically routes supervision for each token between two sources: the teacher and the model itself. The available description does not detail experimental results or specify the conditions under which each source is selected.
The proposal may interest engineers exploring alternatives to standard teacher-based distillation and self-distillation. If you use AI to study or apply the method, avoid entering personal data, confidential code, or other internal data without authorization; use synthetic or anonymized examples where possible.