metaai·lightalo unofficial · independent
Universe / Vision / Sapiens
Vision · 2024

Sapiens

308 keypoints per body, learned from 300 million human images

open source 0.3B / 0.6B / 1B / 2B (Sapiens2 adds 5B) params

Latest: Sapiens (Aug 2024)

Reality Labs' Codec Avatars team's foundation models for seeing people (ECCV 2024): ViTs from 0.3B to 2B parameters pretrained on 300 million human images at native 1024px, delivering state-of-the-art 2D pose estimation, body-part segmentation, depth, and surface-normal prediction — the perceptual groundwork for photoreal telepresence avatars.

Why it matters

Sapiens pretrained ViTs up to 2B parameters on 300 million human images at native 1024px and set state of the art on pose, body-part segmentation, depth and normals — with a 308-keypoint vocabulary dense enough to drive hands and faces. It is the perception groundwork for Reality Labs' photoreal Codec Avatars.

Facts

Try it yourself

Sources

More in Vision

SAM 2ChameleonVideo SealSegment AnythingSAM 3DINOv2

Read the Vision story on the sky →

✦ Open on the map Explore Vision Quiz me