Skip to content
Rota Nacional

Radar ·

ORPO combines SFT and preference optimization

Published on September 19, 2026, the paper introduces ORPO, a method combining preference alignment and supervised fine-tuning without a reference model.

The paper presents ORPO, an approach without a reference model that combines preference alignment with supervised fine-tuning (SFT). During preference-aligned SFT, the method applies a penalty to disfavored generations. The proposal is to assess an alternative to running SFT and preference-alignment stages separately.

The archive description provides no metrics or experimental results. To verify the contribution, consult the original paper under the title “ORPO combines preference optimization with supervised fine-tuning” and examine its methods, comparisons, and results; do not treat the description as evidence of superior performance.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free