Published on October 1, 2026, the text presents multi-teacher on-policy distillation (MOPD): a method that trains a student model on its own completions, using teachers’ per-token log probabilities and a reverse KL objective.
Published on October 1, 2026, the text describes multi-teacher on-policy distillation (MOPD) for language models. A student model is trained on its own completions, using teacher models’ per-token log probabilities and a reverse KL objective.
A configuration called consolidation transfers capabilities from independently trained domain specialists to a student. The technique is presented as a way to combine those capabilities without training the specialists jointly in multi-domain RL. To consult and verify the method, read the original publication and check its definition of the objective, the consolidation setup, and the experimental evidence; the available summary gives no quantitative results.