metaai·lightalo unofficial · independent
Universe / Speech & Sound / Audiobox
Speech & Sound · 2023

Audiobox

Describe the voice and the room — it generates both

closed / product
❝ 184 citationsread 2026-09-03

Latest: Audiobox (Dec 2023, research demos)

Voicebox's successor unifying speech, sound-effect, and soundscape generation with natural-language prompts — 'a young woman speaks with a high pitch, in a large cathedral' — plus voice restyling that keeps the speaker but changes the acoustic scene. Released as interactive research demos with responsible-AI guardrails (audio watermarking, voice authentication) rather than open weights.

Why it matters

Extended Voicebox to sound effects and soundscapes and was the first model to combine voice prompts with text descriptions for freeform voice restyling, surpassing AudioLDM2, VoiceLDM and TANGO on quality. It shipped as a watermarked, research-only demo that Meta retired in February 2026, so today only the paper remains.

Facts

Try it yourself

Lineage

Descends fromVoicebox

See the whole family tree →

Sources

More in Speech & Sound

SeamlessM4T & the Seamless familyMassively Multilingual SpeechAudioCraftNo Language Left BehindUniversal Speech TranslatorSpirit LM

Read the Speech & Sound story on the sky →

✦ Open on the map Explore Speech & Sound Quiz me