Research group guide

Robotic AI & Learning Lab

Sergey Levine

Berkeley RAIL research studies robot learning from demonstrations, reward, prior datasets, and interaction: policies, reinforcement learning, manipulation, and generalization. The independent Levine primer also covers his co-founder work at Physical Intelligence, a separate company.

Official research website ↗

This independent educational primer introduces scientific concepts relevant to research themes associated with Sergey Levine's Robotic AI & Learning Lab at UC Berkeley and discusses publicly documented technology developed by Physical Intelligence. It is not an official UC Berkeley, RAIL Lab, Sergey Levine, or Physical Intelligence course and does not imply endorsement or affiliation.

Questions behind the work

Research questions

01

How can robots acquire behavior from data rather than hand-written rules?

02

How can prior datasets reduce the need for risky new interactions?

03

How can one policy reach many goals across environments?

04

How can generative policies represent different valid actions?

05

How can vision-language knowledge support continuous robot control?

06

How can robots improve from supervised experience while retaining safety and robustness?

Your recommended path

Learn this research

Research primer

How Robots Learn

7 lessons · ~37 minutes

A video-first introduction to robot policies, imitation and reinforcement learning, offline learning, goal-conditioned behavior, generative control, and robot foundation models. Academic RAIL research and Physical Intelligence company research are identified separately.

Watch videos
  1. 01From programming robots to teaching robots
  2. 02Imitation learning versus reinforcement learning
  3. 03How can robots learn without millions of dangerous mistakes?
  4. 04How does a robot learn many goals instead of one task?
  5. 05Generative models for robot control
  6. 06Physical Intelligence: building a foundation model for robots
  7. 07Can robots learn from their own experience?

Key concepts

Behavioral cloning and imitation learning

Explain behavioral cloning and compounding error under distribution shift.

Conservative learning and offline-to-online improvement

Explain conservative value estimates and cautious offline-to-online improvement.

Cross-embodiment training and hierarchical control

Explain cross-embodiment data and the subtask-action-observation hierarchy.

Demonstrations, reward feedback, interventions, and continual improvement

Describe the distinct contributions of demonstrations, interventions, and human reward labels.

Diffusion, flow matching, and continuous robot control

Distinguish diffusion and continuous flow matching from language token prediction.

Generative policies and multimodal actions

Explain why averaging distinct valid actions can be unsafe.

Goal-conditioned policies and representations

Explain how a goal changes the action chosen by one policy.

Learning from autonomous robot experience

Explain how autonomous attempts expose policy-specific failures.

Learning-based control versus hand-engineered robotics

Compare learned visuomotor control with an engineered pipeline.

Offline reinforcement learning and distribution shift

Explain why fixed-data learning cannot test an unfamiliar action.

Planning, subgoals, and action chunking

Explain the roles of subgoals and temporally coherent action chunks.

Reinforcement learning, reward, return, and Q-values

Distinguish a demonstration from reward and explain expected return and Q-values.

Robot policies and end-to-end visuomotor learning

Explain a policy and the closed perception-action loop.

Vision-language-action models and robot foundation models

Distinguish a VLM from a VLA and a company goal from demonstrated capability.

Important papers

2023

ViNT: A Foundation Model for Visual Navigation

Shah D, Sridhar A, Dashora N, Stachowicz K, Black K, Hirose N, Levine S

Why this matters: Berkeley / RAIL research. A visual-navigation foundation model evaluated across robot platforms.

2025

Flow Q-Learning

Park S, Li Q, Levine S

Why this matters: Berkeley / RAIL research. Combines a generative policy with value-guided learning.

2025

π*0.6: a VLA That Learns From Experience

Physical Intelligence, Amin A, Aniceto R, Balakrishna A, Black K, et al

Why this matters: Physical Intelligence research: RECAP combines demonstrations, autonomous attempts, interventions, and human outcome labels; not fully autonomous.

Independent project ideas inspired by this research

Projects you could do

Educational ideas using public or synthetic data. These projects are not offered or supervised by the lab or research group.

Intermediate

Clone a simulated arm

Computational simulation or public-data analysis

Compare an imitation policy on training states and shifted states.

Background
Relevant primer lessons, Basic Python and probability
Data
{"Public simulated demonstration trajectories; no physical robot."}
Output
A short report of errors before and after a controlled distribution shift.

Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.

Intermediate

Measure compounding errors

Computational simulation or public-data analysis

Compare rollout errors with held-out supervised prediction errors.

Background
Relevant primer lessons, Basic Python and probability
Data
{"Synthetic navigation demonstrations and the sl1-2 source map."}
Output
Plots explaining why good one-step predictions do not guarantee stable rollouts.

Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.

Intermediate

Audit an offline benchmark

Computational simulation or public-data analysis

Compare action coverage and optimistic estimates without making safety claims.

Background
Relevant primer lessons, Basic Python and probability
Data
{"A small public offline RL benchmark or a fully synthetic transition dataset."}
Output
A coverage/value report identifying unsupported estimates.

Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.

Intermediate

Reach several simulated goals

Computational simulation or public-data analysis

Condition one navigation policy on different destination inputs.

Background
Relevant primer lessons, Basic Python and probability
Data
{"A toy grid world or public navigation observations."}
Output
A reproducible comparison across familiar and held-out goals.

Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.

Intermediate

Model two valid action modes

Computational simulation or public-data analysis

Compare squared-error regression with a toy diffusion or flow model.

Background
Relevant primer lessons, Basic Python and probability
Data
{"Synthetic left/right navigation action samples."}
Output
A plot showing two valid modes and why their average can collide.

Independent learning idea, not offered or supervised by the lab or Physical Intelligence. Use public data or simulations only; no expensive physical robot, unsafe physical exploration, or real-world deployment.

Related labs and groups

Ranked by shared research topics in the current collection.