Published on September 19, 2026, the paper introduces ORPO, a method combining preference alignment and supervised fine-tuning without a reference model.
The paper presents ORPO, an approach without a reference model that combines preference alignment with supervised fine-tuning (SFT). During preference-aligned SFT, the method applies a penalty to disfavored generations. The proposal is to assess an alternative to running SFT and preference-alignment stages separately.
The archive description provides no metrics or experimental results. To verify the contribution, consult the original paper under the title “ORPO combines preference optimization with supervised fine-tuning” and examine its methods, comparisons, and results; do not treat the description as evidence of superior performance.