An article examines how separate engines can assign inconsistent probabilities to the same trajectories during LLM post-training, and describes TIS and rejecting checkpoint updates.
In reinforcement learning (RL) post-training for language models, different engines for training and inference can assign inconsistent probabilities to the same trajectories. This gap can affect how closely the data used in training corresponds to results evaluated during inference.
The article describes TIS and rejecting checkpoint updates when training and inference rewards diverge. Rejection aims to address this discrepancy, but comes at a cost: it requires additional rollouts. If you use AI to study or apply the material, avoid entering internal or personal data without authorization and review your organization’s data-handling rules.