brains/rt-2 · model record
RT-2
Early vision-language-action model that co-fine-tunes large VLMs (PaLI-X and PaLM-E, up to 55B parameters) on web data and robot trajectories, representing actions as text tokens. Showed transfer of web knowledge to robot control on unseen objects and instructions.
01
Embodiments
robots it has been shown on- Everyday Robots mobile manipulator
02
Papers naming this model
from the papers feedNo paper in the current feed names RT-2 in its title.