Atlas · GenAI 2026
Reinforcement Learning
Learning policies from reward through interaction (Q-learning, policy gradients, PPO); the foundation under RLHF and control.
conceptPeak: 2017Reinforcement LearningAI consensus: 0/3
Prerequisites
MDPs and returns are defined probabilistically.
Policy improvement is an optimization problem.
Recommended reference
Reinforcement Learning: An Introduction (2nd ed.) — Sutton & Barto
Notes from AI deep research
Related skills
- ← is subcategory of: Multi-armed Bandits(3/3)
- → is subcategory of: Machine Learning(3/3)
- ← is subcategory of: GRPO(0/3)