metaai·lightalo unofficial · independent
Universe / Speech & Sound / Massively Multilingual Speech
Speech & Sound · 2023

Massively Multilingual Speech

Massively Multilingual Speech (MMS)

Speech tech for 1,100+ languages

open source superseded 1B (mms-1b-all) params

Latest: MMS 1.0 (May 2023)

Speech-to-text and text-to-speech for 1,107 languages and spoken-language identification for 4,000+ — a 10x leap in coverage, built by pairing wav2vec 2.0 with a surprising data source: recordings of read religious texts (notably the Bible) available in thousands of languages. Halved Whisper's word error rate on its shared languages at the time.

Why it matters

Took open speech recognition from roughly 100 languages to 1,107, with language identification for 4,000+, by pairing wav2vec 2.0 with recordings of read religious texts. It halved Whisper's word error rate on shared languages and made Meta the main supplier of speech tech for low-resource languages.

Facts

Try it yourself

Lineage

Descends fromwav2vec 2.0
Led toSeamlessM4T & the Seamless familyOmnilingual ASR

See the whole family tree →

Sources

More in Speech & Sound

AudioCraftVoiceboxAudioboxNo Language Left BehindUniversal Speech TranslatorSpirit LM

Read the Speech & Sound story on the sky →

✦ Open on the map Explore Speech & Sound Quiz me