01
Research group guide
Robotic AI & Learning Lab
Sergey LevineBerkeley RAIL research studies robot learning from demonstrations, reward, prior datasets, and interaction: policies, reinforcement learning, manipulation, and generalization. The independent Levine primer also covers his co-founder work at Physical Intelligence, a separate company.
Official research website ↗This independent educational primer introduces scientific concepts relevant to research themes associated with Sergey Levine's Robotic AI & Learning Lab at UC Berkeley and discusses publicly documented technology developed by Physical Intelligence. It is not an official UC Berkeley, RAIL Lab, Sergey Levine, or Physical Intelligence course and does not imply endorsement or affiliation.
Questions behind the work
Research questions
02
How can prior datasets reduce the need for risky new interactions?
03
How can one policy reach many goals across environments?
04
How can generative policies represent different valid actions?
05
How can vision-language knowledge support continuous robot control?
06
How can robots improve from supervised experience while retaining safety and robustness?
Your recommended path
Learn this research
Research primer
How Robots Learn
7 lessons · ~37 minutes
A video-first introduction to robot policies, imitation and reinforcement learning, offline learning, goal-conditioned behavior, generative control, and robot foundation models. Academic RAIL research and Physical Intelligence company research are identified separately.
- 01From programming robots to teaching robots
- 02Imitation learning versus reinforcement learning
- 03How can robots learn without millions of dangerous mistakes?
- 04How does a robot learn many goals instead of one task?
- 05Generative models for robot control
- 06Physical Intelligence: building a foundation model for robots
- 07Can robots learn from their own experience?
Key concepts
Behavioral cloning and imitation learning
Explain behavioral cloning and compounding error under distribution shift.
Conservative learning and offline-to-online improvement
Explain conservative value estimates and cautious offline-to-online improvement.
Cross-embodiment training and hierarchical control
Explain cross-embodiment data and the subtask-action-observation hierarchy.
Demonstrations, reward feedback, interventions, and continual improvement
Describe the distinct contributions of demonstrations, interventions, and human reward labels.
Diffusion, flow matching, and continuous robot control
Distinguish diffusion and continuous flow matching from language token prediction.
Generative policies and multimodal actions
Explain why averaging distinct valid actions can be unsafe.
Goal-conditioned policies and representations
Explain how a goal changes the action chosen by one policy.
Learning from autonomous robot experience
Explain how autonomous attempts expose policy-specific failures.
Learning-based control versus hand-engineered robotics
Compare learned visuomotor control with an engineered pipeline.
Offline reinforcement learning and distribution shift
Explain why fixed-data learning cannot test an unfamiliar action.
Planning, subgoals, and action chunking
Explain the roles of subgoals and temporally coherent action chunks.
Reinforcement learning, reward, return, and Q-values
Distinguish a demonstration from reward and explain expected return and Q-values.
Robot policies and end-to-end visuomotor learning
Explain a policy and the closed perception-action loop.
Vision-language-action models and robot foundation models
Distinguish a VLM from a VLA and a company goal from demonstrated capability.
Important papers
2016
End-to-End Training of Deep Visuomotor Policies
Levine S, Finn C, Darrell T, Abbeel P
Why this matters: Berkeley / RAIL research. Joint perception and motor learning for evaluated robot tasks; not universal competence.
2018
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja T, Zhou A, Abbeel P, Levine S
Why this matters: Berkeley / RAIL research. Connects expected return, stochastic policies, and off-policy learning.
2020
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Levine S, Kumar A, Tucker G, Fu J
Why this matters: Berkeley / RAIL research. Explains fixed datasets and the fundamental distribution-shift problem.
2020
Conservative Q-Learning for Offline Reinforcement Learning
Kumar A, Zhou A, Tucker G, Levine S
Why this matters: Berkeley / RAIL research. Introduces conservative values for poorly supported actions.
2020
AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Nair A, Gupta A, Dalal M, Levine S
Why this matters: Berkeley / RAIL research. Links learning from prior datasets to subsequent online improvement; venue remains unconfirmed in the supplied source.
2023
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
Park S, Ghosh D, Eysenbach B, Levine S
Why this matters: Berkeley / RAIL research. Uses latent subgoals to connect a high-level policy to goal-reaching behavior.
2023
ViNT: A Foundation Model for Visual Navigation
Shah D, Sridhar A, Dashora N, Stachowicz K, Black K, Hirose N, Levine S
Why this matters: Berkeley / RAIL research. A visual-navigation foundation model evaluated across robot platforms.
2023
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT)
Zhao TZ, Kumar V, Levine S, Finn C
Why this matters: Berkeley / RAIL research. Multi-institution Berkeley, Stanford, and Meta work: chunking mitigates compounding errors without eliminating them.
2022
Planning with Diffusion for Flexible Behavior Synthesis (Diffuser)
Janner M, Du Y, Tenenbaum JB, Levine S
Why this matters: Berkeley / RAIL research. Berkeley with MIT co-authors: plans trajectories by iterative denoising.
2024
NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration
Sridhar A, Shah D, Glossop C, Levine S
Why this matters: Berkeley / RAIL research. Uses a diffusion policy for goal-directed navigation and exploration.
2025
Flow Q-Learning
Park S, Li Q, Levine S
Why this matters: Berkeley / RAIL research. Combines a generative policy with value-guided learning.
2024
π0: A Vision-Language-Action Flow Model for General Robot Control
Black K, Brown N, Driess D, et al., Levine S, et al
Why this matters: Physical Intelligence research: pretrained vision-language knowledge plus a flow-matching action expert; evaluated tasks do not establish universal control.
2025
π0.5: a Vision-Language-Action Model with Open-World Generalization
Physical Intelligence, Black K, Brown N, Darpinian J, et al
Why this matters: Physical Intelligence research: semantic subtasks and action chunks, with three unseen-home evaluations and explicit limitations.
2025
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Shi LX, Ichter B, Equi M, Ke L, Pertsch K, Vuong Q, Tanner J, Walling A, Wang H, Fusai N, Li-Bell A, Driess D, Groom L, Levine S, Finn C
Why this matters: Physical Intelligence research with Stanford and Berkeley co-authors: hierarchical instruction following; not human-level reasoning.
2025
π*0.6: a VLA That Learns From Experience
Physical Intelligence, Amin A, Aniceto R, Balakrishna A, Black K, et al
Why this matters: Physical Intelligence research: RECAP combines demonstrations, autonomous attempts, interventions, and human outcome labels; not fully autonomous.
2026
π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Physical Intelligence
Why this matters: Physical Intelligence current research: context-steered transfer examples with lower reported success on unseen tasks; transfer remains unresolved.
Independent project ideas inspired by this research
Projects you could do
Educational ideas using public or synthetic data. These projects are not offered or supervised by the lab or research group.
Intermediate
Clone a simulated arm
Computational simulation or public-data analysis
Compare an imitation policy on training states and shifted states.
- Background
- Relevant primer lessons, Basic Python and probability
- Data
- {"Public simulated demonstration trajectories; no physical robot."}
- Output
- A short report of errors before and after a controlled distribution shift.
Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.
Intermediate
Measure compounding errors
Computational simulation or public-data analysis
Compare rollout errors with held-out supervised prediction errors.
- Background
- Relevant primer lessons, Basic Python and probability
- Data
- {"Synthetic navigation demonstrations and the sl1-2 source map."}
- Output
- Plots explaining why good one-step predictions do not guarantee stable rollouts.
Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.
Intermediate
Audit an offline benchmark
Computational simulation or public-data analysis
Compare action coverage and optimistic estimates without making safety claims.
- Background
- Relevant primer lessons, Basic Python and probability
- Data
- {"A small public offline RL benchmark or a fully synthetic transition dataset."}
- Output
- A coverage/value report identifying unsupported estimates.
Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.
Intermediate
Reach several simulated goals
Computational simulation or public-data analysis
Condition one navigation policy on different destination inputs.
- Background
- Relevant primer lessons, Basic Python and probability
- Data
- {"A toy grid world or public navigation observations."}
- Output
- A reproducible comparison across familiar and held-out goals.
Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.
Intermediate
Model two valid action modes
Computational simulation or public-data analysis
Compare squared-error regression with a toy diffusion or flow model.
- Background
- Relevant primer lessons, Basic Python and probability
- Data
- {"Synthetic left/right navigation action samples."}
- Output
- A plot showing two valid modes and why their average can collide.
Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.
Related labs and groups
Ranked by shared research topics in the current collection.
Bronte-Stewart Lab
Helen Bronte-StewartHuman Motor Control and Neuromodulation Laboratory: linking quantitative movement measurements with implanted subthalamic local field potentials and neuromodulation to understand Parkinson’s disease and develop personalized adaptive therapies.
Shared topics: Machine learning, Neural network models
5 lessons · ~27 minutes
Explore the labStephen A. Baccus Lab
Stephen A. BaccusHow retinal circuits transform visual information, adapt and compute; models that connect predictions to mechanisms.
Shared topics: Neural network models
6 lessons · ~45 minutes
Explore the labKarl Deisseroth Lab
Karl DeisserothControlling, mapping and modeling neural circuits across cell types, space and behavior.
Shared topics: Neural network models
9 lessons · ~48 minutes
Explore the lab