humanoidrobots.systems

brains/v-jepa-2 · model record

V-JEPA 2

Self-supervised joint-embedding predictive video model (up to ~1B-parameter ViT-g encoder) pretrained on over 1 million hours of video. The action-conditioned V-JEPA 2-AC, post-trained on under 62 hours of DROID robot video, enables zero-shot pick-and-place planning on Franka arms.

Company
Kind
World model
Access
Open weights
License
MIT
Released
11 Jun 2025
Embodiments
1
01

Embodiments

robots it has been shown on
  • Franka not in the atlas
02

Papers naming this model

from the papers feed
No paper in the current feed names V-JEPA 2 in its title.