How can robots learn without millions of dangerous mistakes?
Explain why fixed-data learning cannot test an unfamiliar action.
Loading video…
# How can robots learn without millions of dangerous mistakes?
*Evidence guide: Offline learning notation is general RL background. CQL, IQL and AWAC are named academic research examples, not Physical Intelligence products.*
Standard online reinforcement learning runs a loop: act, fail, learn, act again. Every improvement needs new interaction with the world. In simulation that is cheap; you can run many trials in parallel. On a real robot, every trial costs time, supervision, and risk.
Here is an alternative. First, collect a dataset: logs from earlier experiments, demonstrations, or other robots doing related things. Then learn the best policy you can from that fixed data, without collecting anything new during training.
This is offline reinforcement learning. In the words of a 2020 tutorial by Levine, Kumar, Tucker and Fu, it uses previously collected data, without additional online data collection. The dataset holds transitions: a state, an action, a reward, and the next state.
Offline RL is not the same as simulation. A simulator lets you try a new action and see what happens. A dataset cannot: it only records what was actually tried. The promise, as the tutorial puts it, is to turn large datasets into powerful decision-making engines.
To improve, the learned policy must choose actions, and the Q-function must score them. For actions the dataset covers, its estimates are anchored to real outcomes. For unfamiliar actions, it is guessing, and some guesses will be too high. The policy, searching for the highest score, finds exactly those errors.
The tutorial calls the fundamental challenge distributional shift: the function is trained under one distribution, but evaluated on a different one. In online RL, the robot would try an over-rated action and discover it was bad. Offline, there is no such correction, so the errors can grow.
One answer is conservatism: be cautious about actions the data does not support. Conservative Q-Learning, or CQL, from Kumar, Levine and colleagues, adds a penalty that pushes down Q-values for actions outside the data. The aim is a Q-function whose value for the learned policy is a lower bound on its true value: if it errs, it errs on the pessimistic side.
A different strategy is to avoid evaluating unseen actions at all. Implicit Q-Learning, from Kostrikov, Nair and Levine, never evaluates actions outside the dataset. It estimates how good the best dataset actions are, then extracts a policy by imitating dataset actions, weighted by how much better than expected they were.
Offline RL does not let a robot safely infer the value of arbitrary actions it has never seen. Conservatism trades away some possible improvement for reliability, and the final policy is limited by what the data covers.
The natural next step is offline-to-online RL. Start from prior data to get useful initial behavior, then practice in the real world, and keep improving. AWAC, from Nair, Gupta, Dalal and Levine, did this on real robots, including a multi-fingered hand and a robot arm opening a drawer.
The broader principle: past experience, plus a limited amount of new experience, can produce a better policy than either alone.
Notation: `(s,a,r,s′)` records a state, action, reward, and next state in a fixed transition dataset.
Sources: [levine2020](https://arxiv.org/abs/2005.01643), [kumar2020](https://papers.nips.cc/paper/2020/hash/0d2b2061826a5df3221116a5085a6052-Abstract.html), [kostrikov2022](https://openreview.net/forum?id=68n2s9ZJWF8), [nair2020](https://arxiv.org/abs/2006.09359).