brains/v-jepa-2 · model record
V-JEPA 2
Self-supervised joint-embedding predictive video model (up to ~1B-parameter ViT-g encoder) pretrained on over 1 million hours of video. The action-conditioned V-JEPA 2-AC, post-trained on under 62 hours of DROID robot video, enables zero-shot pick-and-place planning on Franka arms.
01
Embodiments
robots it has been shown on- Franka
02
Papers naming this model
from the papers feedNo paper in the current feed names V-JEPA 2 in its title.