Socratic LearnCourse overview

Robot Learning & Embodied AI

How Robots Learn

Free viewing — watch in any order. Sign in and enroll if you want quizzes and a certificate.

Generative models for robot control

Explain why averaging distinct valid actions can be unsafe.

Loading video…

# Generative models for robot control *Evidence guide: Generative-model concepts are general background. Diffuser has Berkeley and MIT co-authors; Diffusion Policy is work by Chi and colleagues at Columbia, MIT and TRI, not a Levine paper.* For one situation there may be many good actions. A robot that must get around a chair can go left, or go right. Both are fine. Now train a simple regression policy on demonstrations where people went left half the time and right half the time. Minimizing squared error, the best single prediction is the average: straight ahead, into the chair. Averaging two good answers produced a bad one. The fix is to learn a distribution over good actions, instead of a single average. A generative policy can be sampled: sometimes left, sometimes right, never into the chair. The same kinds of models that generate images can generate actions. Diffusion models are one way to do this. Training takes real examples, adds noise step by step, and trains a network to predict how to remove it. To generate, start from pure noise and repeatedly denoise, until a clean sample appears. In 2022, Janner, Du, Tenenbaum and Levine applied this to whole trajectories. Their model, Diffuser, plans by iteratively denoising a trajectory, so that sampling from the model and planning with it become nearly the same thing. A start and a goal can be fixed while the rest of the plan is denoised. Diffusion policies generate short action sequences from a robot’s observations. Diffusion Policy, by Chi and colleagues at Columbia, MIT and the Toyota Research Institute, showed strong manipulation results. Levine’s group has used diffusion policies too, for example NoMaD, a single policy for both goal-directed navigation and exploration. Flow matching is a close relative. Imagine a simple cloud of random points, and the cloud of useful actions you want. Flow matching learns a velocity field: at every point and every moment, which way should a sample move? Following the arrows from time zero to time one carries random noise into a realistic action. Training is simple to state. Pick a real action and a noise sample. Pick a point partway along the straight path between them, and train the network to predict the direction of that path. At run time, a handful of integration steps produces an action. Now combine two ingredients. Vision-language models, trained on enormous collections of images and text, can describe a scene, answer questions about it, and connect words like mug or sink to what they look like. But that is semantic knowledge. A VLM does not, by itself, know how to move a gripper, or how hard to pull on a towel. A VLA does not merely translate text into motor commands. Language arrives as discrete tokens, but motor commands are continuous numbers that must be produced many times per second. Some VLAs turn actions into tokens; others attach a generative action module, such as diffusion or flow matching. That design choice leads us to Physical Intelligence. Sources: [ho2020](https://arxiv.org/abs/2006.11239), [janner2022](https://arxiv.org/abs/2205.09991), [chi2023](https://arxiv.org/abs/2303.04137), [sridhar2024](https://arxiv.org/abs/2310.07896), [lipman2023](https://arxiv.org/abs/2210.02747), [pi0](https://arxiv.org/abs/2410.24164), [fast2025](https://arxiv.org/abs/2501.09747).