metaai·lightalo unofficial · independent
Universe / World Models & Embodied AI / V-JEPA
World Models & Embodied AI · 2024

V-JEPA

Released the same month as Sora — and betting the opposite way

open source superseded ViT-L 300M / ViT-H 632M params

Latest: V-JEPA (Feb 2024); superseded by V-JEPA 2

Extends JEPA from images to video (February 2024): learns physical-world intuitions by predicting masked spatio-temporal regions of video in representation space, entirely self-supervised from unlabeled video. Achieved strong frozen-evaluation results on motion-centric benchmarks and set the stage for action-conditioned world models.

Why it matters

V-JEPA extended representation prediction from images to video, learning physical intuitions from unlabeled clips with no pixel generation. Released the same month as Sora, it embodied the opposite bet — abstract prediction over pixel synthesis — set strong frozen-evaluation results on motion benchmarks, and set the stage for action-conditioned world models.

Facts

Try it yourself

Lineage

Descends fromI-JEPA
Led toV-JEPA 2

See the whole family tree →

Sources

More in World Models & Embodied AI

PARTNROpenEQAMeta MotivoVL-JEPAMeta Locate 3DEgo-Exo4D

Read the World Models & Embodied AI story on the sky →

✦ Open on the map Explore World Models & Embodied AI Quiz me