Socratic LearnCourse overview

Robot Learning & Embodied AI

How Robots Learn

Free viewing — watch in any order. Sign in and enroll if you want quizzes and a certificate.

Imitation learning versus reinforcement learning

Explain behavioral cloning and compounding error under distribution shift.

Loading video…

# Imitation learning versus reinforcement learning *Evidence guide: Behavioral cloning and RL notation are general background. SAC is Berkeley / RAIL research; DAgger is general background from other researchers.* In imitation learning, a person demonstrates the task. Often they teleoperate the robot: they move a controller, and the robot copies their motion while cameras and joint sensors record everything. Each moment of a demonstration becomes a pair: what the robot observed, and the action the human chose. The simplest method is behavioral cloning. Train the policy to predict the demonstrated action from the observation. It is supervised learning, applied to actions. For continuous actions, a simple version minimizes the squared difference between the demonstrated action and the predicted action. More generally, we maximize the likelihood of the demonstrated actions. But behavioral cloning has a characteristic weakness. The demonstrations only cover situations the human visited. If the policy makes a small error, it drifts into a state slightly unlike anything in the data. There, its prediction is less reliable, the next error is larger, and the robot drifts further. This is compounding error. In ordinary supervised learning, test examples do not depend on the model’s predictions. In control, the robot’s own actions decide what it sees next, so mistakes accumulate over time. Remedies include collecting corrections in the states the policy actually reaches, as in the DAgger algorithm, and predicting short chunks of actions, which we will meet later. Reinforcement learning, or RL, takes a different approach. There is no demonstrated action. The robot is in a state, takes an action, and the environment returns a next state and a reward: a number saying how good the outcome was. For example, plus one when the cup ends up in the sink. The goal is not the next reward but the sum of rewards over time, called the return. G equals r zero, plus gamma times r one, plus gamma squared times r two, and so on. Gamma, the discount factor, is a number slightly below one, so rewards far in the future count a little less. RL searches for a policy that maximizes the expected return. A central tool is the Q-function, Q of s and a. It answers: how good is it to take action a in state s, counting everything that may happen afterward? Conceptually, Q equals the immediate reward, plus the value of what happens next. With a good Q-function, the robot could act by choosing the action with the highest Q. RL is not random motion. Algorithms balance trying new actions with using what already works. Soft actor-critic, from Levine’s group at Berkeley, rewards the policy for staying varied while it succeeds, which helps it explore and learn stably. It is also off-policy: it can learn from experience collected earlier, or by other policies. So the two approaches complement each other. Imitation gives a robot a reasonable starting point. Reinforcement learning offers a way to improve through experience. And a demonstration is not a reward signal: it shows how someone did the task, not how good each outcome was. Modern robot learning increasingly combines both. Notation: `J(π) = E[Σ γ^t r_t]` is expected discounted return; `Q(s,a)` includes the future consequences of an action. Sources: [ross2011](https://proceedings.mlr.press/v15/ross11a), [zhao2023](https://arxiv.org/abs/2304.13705), [haarnoja2018](https://proceedings.mlr.press/v80/haarnoja18b.html).