DPO Direct Preference Optimization
Simpler alternative to RLHF — train directly on preference pairs without a separate reward model.
Stanford 2023. RLHF requires: SFT → reward model → PPO. DPO skips the reward model and PPO. More stable, cheaper, often just as good. Powers much of modern open source finetuning. Llama 3 and later models use DPO or variants.