Voicebox
So good at voices Meta wouldn't release it
Latest: Voicebox (Jun 2023, paper + demos only)
First generative model to solve speech tasks it wasn't explicitly trained for (June 2023): text-guided flow matching enables voice editing, noise removal, style transfer across six languages, and voice cloning from a 2-second sample. Meta deliberately withheld the model and weights, citing voice-impersonation risks — publishing the research with an audio-watermark classifier instead.
Why it matters
Showed one flow-matching model could do zero-shot TTS, editing, denoising and cross-lingual style transfer from a 2-second sample, beating VALL-E on word error rate (1.9% vs 5.9%) while up to 20x faster. Meta withheld the weights over impersonation risk, an early high-profile case of publishing without releasing.
Facts
- A rare and showcase-worthy 'too dangerous to release' decision: 20x faster than prior diffusion-style speech models, could clone a voice from 2 seconds of audio, and Meta shipped a detector-classifier paper instead of the model.
Try it yourself
Official demo page (audio samples) ↗ Read the paper ↗
Lineage
Sources
arXiv ↗Live demo ↗Meta AI blog ↗