metaai·lightalo unofficial · independent
Universe / Speech & Sound / Voicebox
Speech & Sound · 2023

Voicebox

So good at voices Meta wouldn't release it

closed / product

Latest: Voicebox (Jun 2023, paper + demos only)

First generative model to solve speech tasks it wasn't explicitly trained for (June 2023): text-guided flow matching enables voice editing, noise removal, style transfer across six languages, and voice cloning from a 2-second sample. Meta deliberately withheld the model and weights, citing voice-impersonation risks — publishing the research with an audio-watermark classifier instead.

Why it matters

Showed one flow-matching model could do zero-shot TTS, editing, denoising and cross-lingual style transfer from a 2-second sample, beating VALL-E on word error rate (1.9% vs 5.9%) while up to 20x faster. Meta withheld the weights over impersonation risk, an early high-profile case of publishing without releasing.

Facts

Try it yourself

Lineage

Led toAudiobox

See the whole family tree →

Sources

More in Speech & Sound

SeamlessM4T & the Seamless familyMassively Multilingual SpeechAudioCraftNo Language Left BehindUniversal Speech TranslatorSpirit LM

Read the Speech & Sound story on the sky →

✦ Open on the map Explore Speech & Sound Quiz me