V-JEPA 2, a new model that wants to understand the world, by Meta | Meta Comune | Meta WhatsApp | Meta AI Europe | Turtles AI

V-JEPA 2, a new model that wants to understand the world, by Meta
A new video-based self-supervised model integrates visual understanding, prediction and zero-shot robotic planning, leveraging massive Internet datasets and a few hours of robotic data to operate in unfamiliar environments without training
Editorial Team16 June 2025

 

Meta AI presents V‑JEPA2, a self-supervised, video-based world model capable of understanding, anticipating, and planning. Pre-trained at scale, extended to zero-shot robotic capabilities, it demonstrates state-of-the-art performance in visual understanding and robotic control.

Key Points:

  • V‑JEPA2 is a 1.2 billion parameter video-based world model, self-supervised with over 1M hours of internet video.
  • It shows top-1 results on Something‑Somethingv2 (77.3%) and recall-at-5 on Epic‑Kitchens‑100 (39.7%) without language supervision.
  • Using an action-conditioned extension (V‑JEPA2‑AC) trained with only 62h of robotic videos, it plans zero-shot tasks on Franka arms in novel environments.
  • Robotic execution includes reaching, grasping, and pick-and-place with image targets, without rewards or specific training.

The new V-JEPA2 (Video Joint Embedding Predictive Architecture2) is a highly scalable, self-supervised world learning model developed by Meta AI on a 1.2 billion parameter ViT encoder. Primary training leverages over a million hours of video and a million internet images, with latent prediction goals for visual representations, avoiding the pixel-level and favoring predictable semantic structures. This approach has achieved top scores on motion understanding benchmarks (77.3% top-1 on Something-Something v2) and human action anticipation (39.7% recall-at-5 on Epic-Kitchens-100), outperforming task-specific models. Furthermore, after alignment with a large 8 billion parameter language model, V‑JEPA2 achieves excellent results on video‑QA queries: 84.0% on PerceptionTest and 76.9% on TempCompass.

In the next step, the model was extended to action‑conditioned (V‑JEPA2‑AC) by retraining on only 62h of unlabeled robotic videos from the Droid dataset, freezing the encoder and training an autoregressive predictor (~300M parameters) to encode future states as a function of action and current state. V‑JEPA2‑AC was then zero‑shot deployed on Franka arms in two labs, performing reaching, grasping, and pick‑and‑place with image targets, using a CEM-optimized MPC, without any data or environment-specific reward. In benchmark tests, it outperforms methods like Octo and Cosmos, achieving average success rates of 65% on cup grasping and up to 80% on pick-and-place, with faster planning times than Cosmos.

The technical performance comes from a progressive training for spatial/temporal resolution, with ViT-g encoders up to 252K iterations on 64 frames at increasing resolution, and a causal predictor with multi-block attention, able to plan sequences of actions optimized with respect to the target representation. The architectures are released open-source on GitHub and HuggingFace, with checkpoints from ViT-L, H, G.

In addition, the scientific community recognizes V-JEPA2 as a significant step towards video-based world models, capable of concrete use in embodied robotics and other applications such as autonomous vehicles, without the need for extensive annotations or specific training. The model is now available for research and experimentation, with public benchmarks to assess understanding of the physical world from video. V-JEPA2 illustrates how large-scale pre-training combined with modest action-conditioned training can generate models capable of understanding, predicting, and acting in the real world.

In closing, V-JEPA2 paves the way for video-foundation models capable of learning the world through observation and then acting with technical expertise and efficiency, with minimal data specifics as a lever for advanced zero-shot robotic capabilities.

Video