Socratic LearnCourse overview

Robot Learning & Embodied AI

How Robots Learn

Free viewing — watch in any order. Sign in and enroll if you want quizzes and a certificate.

Can robots learn from their own experience?

Explain how autonomous attempts expose policy-specific failures.

Loading video…

# Can robots learn from their own experience? *Evidence guide: Physical Intelligence primary results: RECAP and Hi Robot. Their evaluated improvements do not establish unrestricted autonomy, human-level reasoning, or solved robotics. Remaining problems are current research directions.* Imitation has a hidden limit. A robot can copy a person, but its own mistakes are different. A demonstrator rarely fumbles the way this particular policy does. So demonstration data alone may not contain the experience needed to fix the robot’s failures. In November 2025, Physical Intelligence described pi-star-zero-point-six, a model that learns from experience. It starts from pi-zero-point-six and improves it with a recipe called RECAP: reinforcement learning with experience and corrections, via advantage-conditioned policies. RECAP combines three kinds of data: demonstrations; the robot’s own autonomous attempts; and expert teleoperated interventions, where a person takes over during autonomous execution to correct a mistake. Reward comes from labels. For each episode, people judge whether it succeeded, and the reward is derived from that episode-level success label: a small penalty for every step, and a large penalty if the episode fails. So the value of a situation roughly tracks how many steps remain until success. From these labels, RECAP trains a value function that estimates how close a situation is to success. Comparing the value before and after an action gives its advantage: was this action better or worse than expected? Then comes the key step. The policy is trained on all the data, but each example is tagged with an indicator: was this action an improvement, or not? At run time, the policy is asked for the improving kind of action. This lets a large flow-matching VLA learn from both good and bad experience. Demonstrations teach how people do the task. Autonomous attempts reveal what this particular policy gets wrong. Interventions show how to recover. And reward tells the model which outcomes were better. Autonomous experience here does not mean unrestricted exploration. The authors write that the system is not fully autonomous: it relies on people for reward labels, interventions, and resets, and its exploration is largely greedy. This is supervised practice. Another frontier is following open-ended instructions. Hi Robot, from Physical Intelligence with Stanford and Berkeley co-authors, adds a hierarchy. A high-level vision-language model reads the user’s request and feedback, and chooses a simple command, such as, pick up the bread. A low-level pi-zero policy executes it. Many problems remain open. Data: how can we gather enough diverse physical experience? Generalization: will skills transfer to new objects, homes, and robots? Long horizons: can a robot recover from a mistake eight minutes into a task? Autonomous improvement: can a robot learn safely from its own failures? Evaluation: how do we measure general competence, when a benchmark score does not guarantee real-world reliability? Robustness: what if lighting, friction, or hardware change? And safety: how can learning systems explore and adapt, without dangerous actions? Sources: [pistar06](https://arxiv.org/abs/2511.14759), [hirobot](https://arxiv.org/abs/2502.19417).