Audiobox
Describe the voice and the room — it generates both
Latest: Audiobox (Dec 2023, research demos)
Voicebox's successor unifying speech, sound-effect, and soundscape generation with natural-language prompts — 'a young woman speaks with a high pitch, in a large cathedral' — plus voice restyling that keeps the speaker but changes the acoustic scene. Released as interactive research demos with responsible-AI guardrails (audio watermarking, voice authentication) rather than open weights.
Why it matters
Extended Voicebox to sound effects and soundscapes and was the first model to combine voice prompts with text descriptions for freeform voice restyling, surpassing AudioLDM2, VoiceLDM and TANGO on quality. It shipped as a watermarked, research-only demo that Meta retired in February 2026, so today only the paper remains.
Facts
- Combined a described voice with a described environment for the first time.
- Its demo site included interactive story-soundscape toys; Meta added automatic watermarking to every generated clip.
Try it yourself
Read the paper ↗ Meta AI blog post ↗