Socratic LearnCourse overview

Robot Learning & Embodied AI

How Robots Learn

Free viewing — watch in any order. Sign in and enroll if you want quizzes and a certificate.

How does a robot learn many goals instead of one task?

Explain how a goal changes the action chosen by one policy.

Loading video…

# How does a robot learn many goals instead of one task? *Evidence guide: Goal conditioning and chunking are general concepts. Named navigation and HIQL examples are academic RAIL work; ACT includes Stanford and Meta collaborators.* A robot that has learned to put one block into one box has learned one task. Change the block, the box, or the target, and you may have to start again. How can the same policy learn many behaviors? One answer is to give the policy a goal as an extra input. A goal-conditioned policy, pi of a given s and g, answers three questions at once. Where am I? Where do I want to be? And what should I do next? The goal might be a target position, an image of the desired scene, or, later in this course, a sentence. This turns many separate tasks into one learning problem. Reaching a hundred targets becomes one policy with a hundred possible goals. And every trajectory in a dataset teaches something: wherever the robot ended up, that place can be treated as a goal it successfully reached. Goals raise a question: what is the right internal description of a situation? A raw camera image has hundreds of thousands of numbers, most of them irrelevant, like the texture of the floor or the color of the wall. A useful representation keeps what matters for control: objects, where they are relative to each other, what can be reached, and how long it takes to get somewhere. One idea studied in Levine’s group is to learn representations in which distance means time: two states are close if one can be reached quickly from the other. Such distances can be asymmetric. Going downhill is quick, climbing back is slow. A distance like that is called a quasimetric, and recent papers from Levine and collaborators use quasimetric representations for goal-reaching. Long goals are hard for a single policy. So split them. A high-level planner proposes intermediate subgoals, and a low-level goal-conditioned policy reaches each one. HIQL, from Park, Ghosh, Eysenbach and Levine, learns a high-level policy that predicts a subgoal representation, and a low-level policy that predicts the actions to reach it. A related idea works at the scale of motion. Instead of predicting one tiny motor command at a time, the policy predicts a short sequence of actions, an action chunk, and executes it. In ACT, from Zhao, Kumar, Levine and Finn, predicting chunks reduces the effective horizon of the task by k-fold, mitigating compounding errors. Chunks also make motion more temporally coherent, and a large network can produce many control steps from each query. Chunking also helps reinforcement learning. In 2025, Li, Zhou and Levine ran RL directly over action chunks. The critic evaluates whole chunks, which gives multi-step value backups without off-policy bias, and exploration becomes temporally coherent. Notation: `π(a|s,g)` chooses actions using both the situation and a supplied goal. Sources: [her2017](https://arxiv.org/abs/1707.01495), [myers2025](https://arxiv.org/abs/2509.20478), [zheng2026](https://arxiv.org/abs/2511.07730), [park2023](https://arxiv.org/abs/2307.11949), [zhao2023](https://arxiv.org/abs/2304.13705), [li2025](https://arxiv.org/abs/2507.07969).