V-JEPA
Released the same month as Sora — and betting the opposite way
Latest: V-JEPA (Feb 2024); superseded by V-JEPA 2
Extends JEPA from images to video (February 2024): learns physical-world intuitions by predicting masked spatio-temporal regions of video in representation space, entirely self-supervised from unlabeled video. Achieved strong frozen-evaluation results on motion-centric benchmarks and set the stage for action-conditioned world models.
Why it matters
V-JEPA extended representation prediction from images to video, learning physical intuitions from unlabeled clips with no pixel generation. Released the same month as Sora, it embodied the opposite bet — abstract prediction over pixel synthesis — set strong frozen-evaluation results on motion benchmarks, and set the stage for action-conditioned world models.
Facts
- Released the same month as OpenAI's Sora, it embodied the opposite bet: LeCun argued predicting abstract representations, not generating pixels, is the path to machines that understand the world.
Try it yourself
Code and models on GitHub ↗ Read the paper ↗
Lineage
Sources
GitHub · jepa ↗Meta AI research ↗