01
The reinforcement learning problem
In reinforcement learning an agent interacts with an environment. At each step it observes a state, chooses an action and receives an outcome that may include a reward. The objective is to learn behaviour that maximizes cumulative future reward.
02
Markov decision processes
A Markov decision process models states, actions, transition probabilities, rewards and often a discount factor. The Markov assumption means the current state contains the information needed to predict future dynamics for the purposes of the model.
Gₜ = Rₜ₊₁ + γRₜ₊₂ + γ²Rₜ₊₃ + ...
03
Value functions
A value function estimates expected future return. The optimal value function satisfies the Bellman optimality equation.
V*(s) = maxₐ Σₛ′ P(s′|s,a)[R(s,a,s′) + γV*(s′)]
04
Temporal-difference learning
Temporal-difference methods learn from incomplete episodes by updating estimates toward targets involving successor estimates. This allows an agent to learn continuously from experience rather than waiting for the final outcome of an episode.
05
Q-learning
Q(s,a) ← Q(s,a) + α[r + γ maxₐ′Q(s′,a′) - Q(s,a)]
Q-learning is an off-policy method: the behaviour used to collect data does not necessarily have to match the greedy policy represented by the learned action values.
06
Deep reinforcement learning
Deep reinforcement learning combines reinforcement-learning objectives with neural-network function approximation. This enables much larger state spaces than traditional tabular methods can handle, but introduces instability, sample inefficiency and sensitivity to reward design.