brains/pi06 · model record
π0.6 / π*0.6
π0.6 builds on π0.5; π*0.6 further trains it with RECAP, an advantage-conditioned reinforcement learning method that uses demonstrations, autonomous rollouts and human corrections. Applied to tasks such as espresso making, laundry folding and box assembly.
01
Embodiments
robots it has been shown onNo embodiment is listed in a source yet.
02
Papers naming this model
from the papers feedNo paper in the current feed names π0.6 / π*0.6 in its title.