metaai·lightalo unofficial · independent
Universe / Vision / DINOv3
Vision · 2025

DINOv3

One frozen backbone, from Instagram photos to Mars

open source ViT-S 21M to ViT-7B 6.7B, plus ConvNeXt distillations params

Latest: DINOv3 (Aug 14, 2025)

Meta's flagship self-supervised backbone, released August 14, 2025: a 7B-parameter ViT trained on 1.7B images with no labels, plus distilled ViT-B/ViT-L and ConvNeXt variants, all under a commercial-use license. Produces state-of-the-art dense features for detection, segmentation, and depth — including a dedicated satellite-imagery model trained on 493M Maxar tiles.

Why it matters

DINOv3 scaled self-supervised vision to a 7B ViT trained on 1.7B images with no labels, producing dense features that beat specialist models on segmentation and depth while frozen. A satellite variant cut tree-canopy-height error in Kenya from 4.1m to 1.2m, and it now underpins SAM 3D Body and WRI deforestation monitoring.

Facts

Try it yourself

Lineage

Descends fromDINOv2

See the whole family tree →

Sources

More in Vision

SAM 3SAM 3D (Objects + Body)VGGTPerception Encoder & Perception Language ModelSAM 2Chameleon

Read the Vision story on the sky →

✦ Open on the map Explore Vision Quiz me