01
Imitation Learning
Imitation learning trains a policy from demonstrations instead of requiring an engineer to explicitly specify every rule.
π(a|s) ≈ π_expert(a|s)
Demonstrations may come from humans, existing controllers, expert planners, or other autonomous systems.
02
Reinforcement Learning
Reinforcement learning optimizes behavior through interaction with an environment and a reward signal.
G_t = Σ_{k=0}^{∞} γ^k r_{t+k+1}Robotics introduces additional challenges because physical interaction is costly, dangerous, and slow compared with simulation.
03
Simulation-to-Real Transfer
Policies trained in simulation may fail on hardware because of differences in dynamics, sensors, friction, latency, perception, and environment statistics. Domain randomization, system identification, adaptation, and large real-world datasets are common approaches to this gap.