Language Technology Partner Program + BOUQuET
Partners contribute underserved languages; the resulting models go open
Latest: Announced Feb 7, 2025 (with UNESCO)
Program sourcing speech recordings, transcriptions, and translations for underserved languages from partners — governments, communities, researchers — in support of UNESCO's Decade of Indigenous Languages; resulting models are open-sourced. Announced alongside BOUQuET, an open multilingual translation benchmark. The Government of Nunavut joined early, contributing Inuktitut and Inuinnaqtun data.
Why it matters
Turns Meta's low-resource language work into a data supply chain: partners contribute 10+ hours of transcribed speech and translated sentences, resulting models are open-sourced, and the effort supports UNESCO's Decade of Indigenous Languages. BOUQuET adds a paragraph-level, multi-way translation benchmark spanning 275 language varieties under CC-BY-4.0.
Facts
- Partners commit 10+ hours of transcribed speech and 200+ translated sentences per language.
- The Omnilingual ASR Corpus (Nov 2025), spanning 350 underserved languages, was curated with these global partners.
Try it yourself
BOUQuET dataset on Hugging Face ↗ Read the BOUQuET paper (EMNLP 2025) ↗ Program announcement (Meta Newsroom) ↗