Physical Intelligence: building a foundation model for robots
Distinguish a VLM from a VLA and a company goal from demonstrated capability.
Loading video…
# Physical Intelligence: building a foundation model for robots
*Evidence guide: Physical Intelligence is a separate company. Its mission is a company-stated goal; π0 and π0.5 results are limited demonstrated capabilities in the reported evaluations.*
This video covers research from Physical Intelligence, a robotics AI company founded in 2024. It is a separate organization from Berkeley’s RAIL lab. Sergey Levine is one of its co-founders, along with researchers including Karol Hausman, Chelsea Finn, and Brian Ichter.
The company describes its mission as bringing general-purpose AI into the physical world, by developing foundation models and learning algorithms for robots. Its papers describe generalist robot policies: one model, trained on many robots and tasks, that can be prompted or fine-tuned for new ones. That is a stated goal, not a demonstrated capability.
The analogy is to language models. A large language model maps text to text. A robot foundation model maps images, language, and the robot’s state to actions. But robotics is not language modeling with motors. Actions must be continuous, fast, and physically safe.
Cross-embodiment learning trains one policy on data from different robot bodies: a single arm, two arms, an arm on a mobile base. Their joints, cameras and action spaces differ. The hope is that shared structure, like what grasping a mug involves, transfers anyway. How far it transfers is still an open research question.
Physical Intelligence’s first generalist policy, pi-zero, was released in October 2024. The paper describes it as a vision-language-action flow model for general robot control.
Pi-zero starts from PaliGemma, a pretrained vision-language model with about three billion parameters, which brings visual and semantic knowledge. To it, the authors add an action expert of about three hundred million parameters that produces continuous robot actions. In total, about three point three billion parameters.
The action expert uses flow matching. Language models predict discrete tokens one at a time. Motor commands are continuous, and must be generated quickly. Pi-zero generates a chunk of fifty future actions at once, and can control robots at up to fifty times per second.
As with language models, training has two stages. Pre-training on the broad mixture teaches general physical skills. Post-training on a smaller, high-quality dataset specializes the model for a particular task.
The next model, pi-zero-point-five, released in April 2025, targets a different problem: open-world generalization. Can a robot work in a home it has never seen?
Pi-zero-point-five uses one model at two levels. Given a goal like, clean the kitchen, it first predicts a high-level semantic subtask in words, such as, pick up the plate. Then its flow-matching action expert produces a chunk of low-level motor actions for that subtask. After the chunk, a new observation arrives, and the model picks the next subtask.
It was evaluated in three real homes that were not in the training data, on tasks like putting dishes in the sink, making a bed, and cleaning up a bedroom floor, including multi-stage cleanups lasting ten to fifteen minutes.
The authors are explicit about limits. The goal was generalization, not new skills or high dexterity. The robot struggled with unfamiliar drawer handles and hard-to-open cabinets, sometimes got distracted while choosing subtasks, and handled relatively simple prompts. It is not a robot that can clean any home.
Sources: [berkeley-vcr](https://vcresearch.berkeley.edu/node/28965), [cofounders](https://ai.stanford.edu/~cbfinn/), [pi-home](https://www.pi.website/), [pi0](https://arxiv.org/abs/2410.24164), [oxe2024](https://arxiv.org/abs/2310.08864), [pi05](https://arxiv.org/abs/2504.16054).