Atlas · GenAI 2026
Direct Preference Optimization
DPO / KTO / ORPO
conceptPeak: 2024AlignmentAI consensus: 3/3
Prerequisites
- mediumRLHF
DPO was created to simplify RLHF — understanding what RLHF does helps understand what DPO replaces and why
- hardLLM Fine-Tuning
DPO modifies the SFT loss function with preference pairs — SFT is the computational foundation
Recommended reference
Rafailov et al. (2023) 'Direct Preference Optimization: Your Language Model is Secretly a Reward Model' — NeurIPS; elegantly eliminates the reward model from RLHF
Notes from AI deep research
Anthropic Opus
Rafailov (2023). Eliminuje reward model. Prostszy, tanszy, szybszy niz RLHF. Default alignment
OpenAI Deep Research
DPO upraszcza pipeline [OA#17]
Google Deep Think
Bez pośrednich modeli nagrody [G#64]
Related skills
- → is subcategory of: LLM Fine-Tuning(3/3)
- → is an instance of: RLHF(1/3)