metaai·lightalo unofficial · independent
Universe / Vision / SAM 2
Vision · 2024

SAM 2

One click on one frame tracks an object through a whole video

open source superseded Hiera tiny 39M / small 46M / base+ 81M / large 224M params

Latest: SAM 2.1 (Sep 2024), plus the SAM 2.1 Developer Suite with open training code

Extends promptable segmentation to video with a streaming-memory transformer: prompt an object once and SAM 2 tracks its 'masklet' across frames in real time (~44 fps), surviving occlusions and reappearances. Also 6x faster and more accurate than SAM on still images. Apache 2.0 code and weights, July 2024.

Why it matters

SAM 2 extended promptable segmentation to video: one click on one frame tracks an object through occlusions at about 44 fps, and it is 6x faster than SAM on stills. Its SA-V dataset (51,000 videos, 643,000 masklets) was roughly 50x larger than prior video segmentation data, making it the standard video-mask tool.

Facts

Try it yourself

Lineage

Descends fromSegment Anything
Led toSAM 3

See the whole family tree →

Sources

More in Vision

ChameleonSapiensVideo SealDINOv2SAM 3D (Objects + Body)DINOv3

Read the Vision story on the sky →

✦ Open on the map Explore Vision Quiz me