SeamlessM4T & the Seamless family
A real-life Babel Fish that keeps your voice
Latest: SeamlessM4T v2 + SeamlessExpressive/Streaming (Nov 2023); Nature paper Jan 2025
One model for speech-to-speech, speech-to-text, text-to-speech, text-to-text translation and ASR across ~100 languages (Aug 2023). The v2 suite (Nov 2023) added SeamlessExpressive — preserving your tone, pauses, and emotion across languages — and SeamlessStreaming, translating with ~2-second latency before the speaker finishes. Published in Nature in January 2025.
Why it matters
Collapsed speech-to-speech, speech-to-text, text-to-speech and text translation into one model for roughly 100 languages, then added expressive translation that preserves tone and pauses and streaming translation at about two seconds of latency. Published in Nature in January 2025, it is the reference open system for speech translation.
Facts
- A real-life Babel Fish — speak in one of ~100 languages, hear it in another, published on the pages of Nature.
- The seamless_communication repo has ~11,900 stars.
- SeamlessExpressive preserves your speech rate, pauses, emotion and vocal style across languages.
- SeamlessAlign's 470,000 hours is the largest open multimodal translation corpus ever mined — about 53 years of continuous audio.
- 68 credited contributors ('Seamless Communication' team).
Try it yourself
Model on Hugging Face (SeamlessM4T v2 Large) ↗ Run it with Transformers (docs) ↗ Meta's Seamless demo site ↗
Lineage
Sources
arXiv · 2308.11596 ↗arXiv · 2312.05187 ↗Nature · d41586 025 00497 2 ↗Nature · s41586 024 08359 z ↗GitHub · seamless_communication ↗Meta AI blog ↗