Read the paper

Official title: V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Meta’s predictive-representation approach to video understanding and action-conditioned world models for robot planning.