siddhant

Knowledge / Machine Learning

Reinforcement Learning

Learning decisions through interaction, rewards and long-term consequences.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

The reinforcement learning problem

In reinforcement learning an agent interacts with an environment. At each step it observes a state, chooses an action and receives an outcome that may include a reward. The objective is to learn behaviour that maximizes cumulative future reward.

02

Markov decision processes

A Markov decision process models states, actions, transition probabilities, rewards and often a discount factor. The Markov assumption means the current state contains the information needed to predict future dynamics for the purposes of the model.

Gₜ = Rₜ₊₁ + γRₜ₊₂ + γ²Rₜ₊₃ + ...

03

Value functions

A value function estimates expected future return. The optimal value function satisfies the Bellman optimality equation.

V*(s) = maxₐ Σₛ′ P(s′|s,a)[R(s,a,s′) + γV*(s′)]

04

Temporal-difference learning

Temporal-difference methods learn from incomplete episodes by updating estimates toward targets involving successor estimates. This allows an agent to learn continuously from experience rather than waiting for the final outcome of an episode.

05

Q-learning

Q(s,a) ← Q(s,a) + α[r + γ maxₐ′Q(s′,a′) - Q(s,a)]

Q-learning is an off-policy method: the behaviour used to collect data does not necessarily have to match the greedy policy represented by the learned action values.

06

Deep reinforcement learning

Deep reinforcement learning combines reinforcement-learning objectives with neural-network function approximation. This enables much larger state spaces than traditional tabular methods can handle, but introduces instability, sample inefficiency and sensitivity to reward design.

References

  1. Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction, 2nd edition.
    https://incompleteideas.net/book/the-book-2nd.html
  2. Stanford University, CS229 Machine Learning. Course materials covering supervised and unsupervised learning, learning theory, regularization, SVMs and reinforcement learning.
    https://cs229.stanford.edu/

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.