VL-JEPA
Predict the embedding, skip the tokens: JEPA outlives its champion
Latest: VL-JEPA (Dec 2025)
A vision-language JEPA (paper December 11, 2025): instead of autoregressively generating tokens like classic VLMs, it predicts continuous embeddings of target text. At only 1.6B parameters it surpasses CLIP, SigLIP2, and Meta's own Perception Encoder across eight video classification and eight retrieval benchmarks while matching classical VLMs on VQA.
Why it matters
VL-JEPA showed JEPA's predict-the-embedding idea works for vision-language: at 1.6B parameters it beats CLIP, SigLIP2 and Meta's own Perception Encoder on video classification and retrieval while matching token-generating VLMs on VQA with 50% fewer trainable parameters. Published the month after LeCun's exit, it proved the JEPA program outlived its champion.
Facts
- Achieves its results with 50% fewer trainable parameters than equivalent token-space VLM training — published the month after Yann LeCun's departure was announced, proving the JEPA program outlives its champion at Meta.