Course · PhD & beyond · 117 pp
Reinforcement Learning
From the Bellman Equation to Language Models
Read-only for now. A complete course in sequential decision making: dynamic programming, temporal-difference learning, policy gradients, deep RL, offline methods, and policy optimisation for large language models.