Socratic LearnAll courses

Robot Learning & Embodied AI

How Robots Learn

A video-first introduction to robot policies, imitation and reinforcement learning, offline learning, goal-conditioned behavior, generative control, and robot foundation models. Academic RAIL research and Physical Intelligence company research are identified separately.

3 modules · 7 lessons · 1h · mastery threshold 80

Watch videos free — no sign-in

About this course

Background for Research in Sergey Levine's RAIL Lab and Physical Intelligence This independent educational primer introduces scientific concepts relevant to research themes associated with Sergey Levine's Robotic AI & Learning Lab at UC Berkeley and discusses publicly documented technology developed by Physical Intelligence. It is not an official UC Berkeley, RAIL Lab, Sergey Levine, or Physical Intelligence course and does not imply endorsement or affiliation. For ambitious high-school students and early undergraduates with algebra and basic probability. Approximately 50 minutes including seven video-first lessons and assessments. Mastery-only certificate eligibility at 80%, without a capstone or Research Defense.

Syllabus

Module 1

How Robots Learn from Experience

Policies, demonstrations, rewards, and learning safely from prior data.

Module 2

From Individual Skills to General Robot Behavior

Goals, useful representations, action chunks, and generative policies.

Module 3

Robot Foundation Models and Physical Intelligence

Levine's Physical Intelligence work: vision-language-action models, hierarchical control, and improvement from supervised real-world experience.

Concepts you'll master

  • Behavioral cloning and imitation learning

    Explain behavioral cloning and compounding error under distribution shift.

  • Conservative learning and offline-to-online improvement

    Explain conservative value estimates and cautious offline-to-online improvement.

  • Cross-embodiment training and hierarchical control

    Explain cross-embodiment data and the subtask-action-observation hierarchy.

  • Demonstrations, reward feedback, interventions, and continual improvement

    Describe the distinct contributions of demonstrations, interventions, and human reward labels.

  • Diffusion, flow matching, and continuous robot control

    Distinguish diffusion and continuous flow matching from language token prediction.

  • Generative policies and multimodal actions

    Explain why averaging distinct valid actions can be unsafe.

  • Goal-conditioned policies and representations

    Explain how a goal changes the action chosen by one policy.

  • Learning from autonomous robot experience

    Explain how autonomous attempts expose policy-specific failures.

  • Learning-based control versus hand-engineered robotics

    Compare learned visuomotor control with an engineered pipeline.

  • Offline reinforcement learning and distribution shift

    Explain why fixed-data learning cannot test an unfamiliar action.

  • Planning, subgoals, and action chunking

    Explain the roles of subgoals and temporally coherent action chunks.

  • Reinforcement learning, reward, return, and Q-values

    Distinguish a demonstration from reward and explain expected return and Q-values.

  • Robot policies and end-to-end visuomotor learning

    Explain a policy and the closed perception-action loop.

  • Vision-language-action models and robot foundation models

    Distinguish a VLM from a VLA and a company goal from demonstrated capability.