Sapiens
308 keypoints per body, learned from 300 million human images
Latest: Sapiens (Aug 2024)
Reality Labs' Codec Avatars team's foundation models for seeing people (ECCV 2024): ViTs from 0.3B to 2B parameters pretrained on 300 million human images at native 1024px, delivering state-of-the-art 2D pose estimation, body-part segmentation, depth, and surface-normal prediction — the perceptual groundwork for photoreal telepresence avatars.
Why it matters
Sapiens pretrained ViTs up to 2B parameters on 300 million human images at native 1024px and set state of the art on pose, body-part segmentation, depth and normals — with a 308-keypoint vocabulary dense enough to drive hands and faces. It is the perception groundwork for Reality Labs' photoreal Codec Avatars.
Facts
- Its 308-keypoint pose vocabulary is far denser than standard 17-keypoint benchmarks — detailed enough to drive lifelike avatar hands and faces.
- 5.4k GitHub stars.
- Beat prior SOTA by 7.6 mAP on pose and 17.1 mIoU on segmentation.
- Sapiens2's 5B model is reported as the highest-FLOPs vision transformer to date (~15.7 TFLOPs per inference).
- The lineage feeds Codec Avatars: the large-scale avatar pretraining work shares authors (e.g.
Try it yourself
Sapiens models on Hugging Face ↗ sapiens-pose-1b on Hugging Face ↗ Read the paper ↗
Sources
arXiv ↗GitHub · sapiens ↗Hugging Face ↗meta.com ↗rawalkhirodkar.github.io ↗