V-JEPA 2
It learned physics from video — then drove a robot arm it had never met
Latest: V-JEPA 2.1 (Mar 16, 2026) — new recipe for high-quality, temporally consistent dense features; family now spans 80M to 2B parameters
Meta's flagship world model, released June 11, 2025: trained on over 1 million hours of video plus only ~62 hours of robot data, its action-conditioned variant (V-JEPA 2-AC) enabled zero-shot planning on real Franka robot arms in unseen labs. Meta simultaneously released three physical-reasoning benchmarks (IntPhys 2, MVPBench, CausalVQA).
Why it matters
V-JEPA 2 is the strongest evidence yet for JEPA-style world models: trained on over a million hours of video plus under 62 hours of robot data, its action-conditioned variant planned zero-shot on Franka arms in labs it never saw — 100% on reaching, 80% on pick-and-place — and set new action-anticipation records.
Facts
- It learned physics by watching YouTube-scale video — then picked up objects with a robot arm it had never controlled.
- Zero-shot robot results: 100% success on reaching, 80% on cup pick-and-place, on robots the model never trained on.
- Planning ~15x faster than NVIDIA's Cosmos world model per the paper (Meta's press materials said up to 30x).
- V-JEPA 2 also jumped the EK100 action-anticipation record from 27.6 to 39.7 recall@5.
- Trained on 1M+ hours of video plus under 62 hours of unlabeled robot footage, then deployed zero-shot on robot arms in two different labs 'without collecting any data from the robots in these environments.' Also set records on Epic-Kitchens-100 action…
Try it yourself
vjepa2-vitl on Hugging Face ↗ Run the official demo notebook in Colab ↗ Full V-JEPA 2 collection on Hugging Face ↗
Lineage
Sources
arXiv ↗GitHub · vjepa2 ↗Meta AI blog ↗