Skip to content
Rota Nacional

Radar ·

On-policy self-distillation for reasoning in LLMs

Published on September 28, 2026, the material presents Self-Distilled Reasoner, an on-policy self-distillation approach for language models. Its preview describes student-sampled trajectories and dense teacher supervision at token level.

The paper presents Self-Distilled Reasoner, an on-policy self-distillation approach for LLMs. According to the preview, training uses trajectories sampled by the student, while the teacher provides dense supervision at token level. The material is dated September 28, 2026.

The proposal addresses distribution mismatch between training and inference in off-policy distillation, an issue relevant to reasoning systems built with LLMs. The available description does not detail experimental results. To assess its claims and methods, consult the original paper and check its methodology, experiments, and limitations; if using AI to study it, avoid entering sensitive organizational data.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free