metaai·lightalo unofficial · independent
Universe / Speech & Sound / wav2vec 2.0
Speech & Sound · 2020

wav2vec 2.0

Speech recognition from raw audio, almost no labels

open source superseded 95M (BASE) / 317M (LARGE) params

Latest: wav2vec 2.0 (2020); descendants XLS-R, MMS, Omnilingual w2v 7B (2025)

Self-supervised speech representation learning (NeurIPS 2020): the model learns from raw unlabeled audio, then needs astonishingly little labeled data — 10 minutes of transcriptions plus 53K hours of unlabeled speech beat prior systems trained on 100x more labels. The foundation under MMS, Seamless, and 2025's Omnilingual ASR.

Why it matters

Showed speech recognition could learn mostly from unlabeled audio: 53K unlabeled hours plus only 10 minutes of transcripts beat systems trained on 100x more labels. That recipe became the backbone of Meta's MMS, Seamless and Omnilingual ASR, and of most self-supervised speech research since.

Facts

Try it yourself

Lineage

Led toMassively Multilingual SpeechUniversal Speech Translator

See the whole family tree →

Sources

More in Speech & Sound

No Language Left BehindSeamlessM4T & the Seamless familyAudioCraftVoiceboxAudioboxSpirit LM

Read the Speech & Sound story on the sky →

✦ Open on the map Explore Speech & Sound Quiz me