MDL-1666EST.2025 · IDX.090
Robotics AIPilot deployment

π*0.6 (pi-star-0.6)

Physical Intelligence · USA · 2025

The π model that learns from its own mistakes — an espresso shift that lasts all day.

wujec.ai score

8.6/10

Community score

no votes yet
Sign in to rate

π*0.6, presented by Physical Intelligence on 17 November 2025, addresses the structural weakness of robot policies trained purely by imitation: a small early error pushes the robot into states no demonstration ever covered, and the mistakes compound. The answer is RECAP — reinforcement learning with experience and corrections via advantage-conditioned policies — which layers three sources of learning on top of one another. The model starts from human demonstrations, then absorbs real-time corrections made by an expert when it goes wrong, and finally improves on its own from autonomous trials, with a value function telling it which of its own attempts were actually worth imitating. The demonstrations were chosen to test endurance rather than novelty. The robot pulled espresso drinks continuously for an 18-hour stretch, folded 50 unfamiliar laundry items in a home it had not worked in before, and assembled and labelled 59 real cardboard boxes in a factory setting. Against the imitation-trained baseline, Physical Intelligence reports roughly double the throughput and less than half the failure rate on the hardest tasks — the company summarised the result as evidence that reinforcement learning is back on the table for real-world manipulation. π*0.6 is a closed model; Physical Intelligence published a model card and a technical report but no weights.

#VLA model#reinforcement learning#RECAP#self-improvement#closed weights
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review