From programming robots to teaching robots
Explain a policy and the closed perception-action loop.
Loading video…
# From programming robots to teaching robots
*Evidence guide: General background and teaching examples introduce policies. The named 2016 visuomotor experiment is Berkeley research, not a Physical Intelligence result.*
Picture a factory robot arm welding a car door. It repeats the same motion thousands of times a day, with great precision. Nothing about it is learned. Engineers measured the parts, planned the path, and wrote the program.
The classical recipe is a pipeline. Engineer the perception: detect the part. Estimate the state: where exactly is it? Plan a trajectory. Write a controller to follow it. Then test, find the failures, and fix them by hand. In a factory this works beautifully, because the world has been engineered to stay the same.
So instead of writing every rule, robot learning asks a different question. Can the robot acquire behavior from experience? Data goes into a learning algorithm, which trains a neural-network policy, and the policy produces behavior. To change the behavior, you change the data or the objective, not thousands of hand-written rules.
Three words will appear in every video. The first is observation: what the robot senses right now. That can be a camera image, the angles of its joints, readings from force or touch sensors, and, in modern systems, a language instruction, such as: put the cup in the sink.
The second is action: the command sent to the motors. It might be target joint positions, joint velocities, or a desired movement of the gripper, sent many times per second.
The third is the policy, written pi of a given s. Read it as a rule that answers one question: given the current situation, what should the robot do next? Here s stands for the state, or what the robot observes, and a for the action. A policy is not a pre-recorded path. It reacts: if the cup slides, the next observation changes, and so does the action.
Put these together and you get a loop. The robot perceives, the policy chooses an action, the action changes the environment, and a new observation arrives. Learning sits on top of this loop, using what happened to improve the policy.
One influential idea is end-to-end learning. Instead of separate hand-built modules for vision and control, a single neural network maps camera pixels directly to motor commands, and both parts are trained together for the task.
In 2016, Levine, Finn, Darrell and Abbeel trained convolutional networks with about ninety-two thousand parameters to map raw camera images directly to the torques at a robot’s motors. The robot learned tasks like hanging a coat hanger on a rack, inserting a block into a shape-sorting cube, and screwing a cap onto a bottle. The point: perception could be shaped by what the task needs, instead of being designed separately.
Throughout this course, three bottlenecks keep returning. Data, because robot experience is slow and expensive to collect. Generalization, because a skill must work with new objects and places. And experience, because a robot must eventually improve from its own attempts.
Notation: `a ~ π(a|s)` means sample an action from the policy conditioned on the current situation.
Sources: [levine2016](https://jmlr.org/papers/v17/15-522.html).