metaai·lightalo unofficial · independent
Universe / Vision / Perception Encoder & Perception Language Model
Vision · 2025

Perception Encoder & Perception Language Model

A vision encoder whose best features hide in its middle layers

open source PE Core: B/16 90M, L/14 320M, G/14 1.88B; PLM: 1B / 3B / 8B params

Latest: PE + PLM (Apr 2025)

FAIR's April 2025 open perception stack: Perception Encoder (PE) is a large-scale vision encoder whose intermediate layers surprisingly hold the best embeddings, topping CLIP-style models on image and video tasks; the Perception Language Model (PLM, 1B/3B/8B) is a fully open, reproducible VLM released with PLM-VideoBench for fine-grained video understanding.

Why it matters

Perception Encoder found that a CLIP-style encoder's best embeddings hide in its intermediate layers, not its output, and used that to top CLIP-class models on image and video tasks. PLM paired it with a fully open, non-distilled VLM and 2.5M new human-labeled video QA samples — a deliberate reproducibility stance against proprietary distillation.

Facts

Try it yourself

Sources

More in Vision

SAM 3SAM 3D (Objects + Body)DINOv3VGGTSAM 2Chameleon

Read the Vision story on the sky →

✦ Open on the map Explore Vision Quiz me