Published on September 28, 2026, the material presents Self-Distilled Reasoner, an on-policy self-distillation approach for language models. Its preview describes student-sampled trajectories and dense teacher supervision at token level.
The paper presents Self-Distilled Reasoner, an on-policy self-distillation approach for LLMs. According to the preview, training uses trajectories sampled by the student, while the teacher provides dense supervision at token level. The material is dated September 28, 2026.
The proposal addresses distribution mismatch between training and inference in off-policy distillation, an issue relevant to reasoning systems built with LLMs. The available description does not detail experimental results. To assess its claims and methods, consult the original paper and check its methodology, experiments, and limitations; if using AI to study it, avoid entering sensitive organizational data.