Published on December 15, 2025, the post describes an Olmo 3 training pipeline: SFT on curated reasoning traces, followed by preference tuning with DPO and RLVR using Group Relative Policy Optimization.
Published on December 15, 2025, the post presents the Olmo 3 training pipeline. It starts with SFT on curated reasoning traces, then applies preference tuning with DPO and RLVR using Group Relative Policy Optimization.
According to the post, combining SFT and DPO is a better starting point for RL than using SFT alone or applying RL directly to the base model. The publication highlights the role of data curation and preference tuning. To verify the account, consult the original post and compare its descriptions of the methods and sequence; the supplied material gives no quantitative results.