IT lexicon AI & ML DPO

DPO Direct Preference Optimization

AI & ML På svenska → Updated: 2026-05-23

Simpler alternative to RLHF — train directly on preference pairs without a separate reward model.

Stanford 2023. RLHF requires: SFT → reward model → PPO. DPO skips the reward model and PPO. More stable, cheaper, often just as good. Powers much of modern open source finetuning. Llama 3 and later models use DPO or variants.

← Back to the lexicon