A paper published on July 11, 2026 combines environmental rewards with self-distillation and reports higher average accuracy on five multimodal benchmarks.
Published on July 11, 2026, the paper on RLSD combines environmental rewards with a self-distillation signal to adjust updates per token along a training trajectory. The proposal differs from the sequence-wide uniform advantage used by GRPO.
On Qwen3-VL-8B, the work reports higher average accuracy than the base model and GRPO on five multimodal reasoning benchmarks. To check the results, consult the original paper and verify its experimental setup, benchmarks, and reported metrics; the available summary does not specify those details.