Atlas · GenAI 2026

Direct Preference Optimization

DPO / KTO / ORPO

conceptPeak: 2024AlignmentAI consensus: 3/3

Prerequisites

  • mediumRLHF

    DPO was created to simplify RLHF — understanding what RLHF does helps understand what DPO replaces and why

  • DPO modifies the SFT loss function with preference pairs — SFT is the computational foundation

Recommended reference

Rafailov et al. (2023) 'Direct Preference Optimization: Your Language Model is Secretly a Reward Model' — NeurIPS; elegantly eliminates the reward model from RLHF

Notes from AI deep research

Anthropic Opus

Rafailov (2023). Eliminuje reward model. Prostszy, tanszy, szybszy niz RLHF. Default alignment

OpenAI Deep Research

DPO upraszcza pipeline [OA#17]

Google Deep Think

Bez pośrednich modeli nagrody [G#64]

Related skills