metaai·lightalo unofficial · independent
Universe / World Models & Embodied AI / VL-JEPA
World Models & Embodied AI · 2025

VL-JEPA

Predict the embedding, skip the tokens: JEPA outlives its champion

closed / product 1.6B params
❝ 54 citationsread 2026-09-03

Latest: VL-JEPA (Dec 2025)

A vision-language JEPA (paper December 11, 2025): instead of autoregressively generating tokens like classic VLMs, it predicts continuous embeddings of target text. At only 1.6B parameters it surpasses CLIP, SigLIP2, and Meta's own Perception Encoder across eight video classification and eight retrieval benchmarks while matching classical VLMs on VQA.

Why it matters

VL-JEPA showed JEPA's predict-the-embedding idea works for vision-language: at 1.6B parameters it beats CLIP, SigLIP2 and Meta's own Perception Encoder on video classification and retrieval while matching token-generating VLMs on VQA with 50% fewer trainable parameters. Published the month after LeCun's exit, it proved the JEPA program outlived its champion.

Facts

Try it yourself

Lineage

Descends fromV-JEPA 2

See the whole family tree →

Sources

More in World Models & Embodied AI

Meta Locate 3DV-JEPAPARTNROpenEQAMeta MotivoI-JEPA

Read the World Models & Embodied AI story on the sky →

✦ Open on the map Explore World Models & Embodied AI Quiz me