metaai·lightalo unofficial · independent
Universe / Speech & Sound / SAM Audio
Speech & Sound · 2025

SAM Audio

SAM Audio (Segment Anything in Audio)

Segment Anything — for sound

open source

Latest: SAM Audio v1 (Dec 16, 2025) — checkpoints small/base/large plus -tv variants; no 1.1 or newer release as of Sep 2026 (no tagged GitHub releases; arXiv paper still v1)

Family of open-weight foundation models that isolate any sound from a complex mixture using text prompts, visual prompts (click the object making the sound in a video), or time-span prompts. A flow-matching diffusion transformer built on Meta's Perception Encoder Audiovisual (PE-AV), it achieves state-of-the-art results across speech, music, instrument, and general sound separation.

Why it matters

Brings the Segment Anything idea to sound: describe a source, click the object making it in a video, or mark a time span, and it isolates that sound from a mixture with state-of-the-art results across speech, music and general audio. Open weights and a browser playground make source separation usable by editors, not only researchers.

Facts

Try it yourself

Lineage

Descends fromSAM 3

See the whole family tree →

Sources

More in Speech & Sound

Omnilingual ASRLanguage Technology Partner Program + BOUQuETSpirit LMSeamlessM4T & the Seamless familyMassively Multilingual SpeechAudioCraft

Read the Speech & Sound story on the sky →

✦ Open on the map Explore Speech & Sound Quiz me