Spirit LM
Text and speech in one token stream — with the emotion intact
Latest: Spirit LM Base & Expressive 7B (weights Oct 18, 2024)
Meta's first open multimodal language model that freely mixes text and speech in one token stream (paper Feb 2024, weights Oct 2024). The Expressive variant adds pitch and style tokens so it can carry emotion — anger, surprise, excitement — across modalities, enabling speech that actually sounds like it means it. FAIR Noncommercial Research License.
Why it matters
Meta's first open model that interleaves spoken and written tokens in one stream, so it can continue a text prompt in speech or the reverse, and its Expressive variant carries pitch and style across modalities. It is a reference design for speech-native language models, released under a non-commercial research license.
Facts
- Trained by interleaving text and speech tokens word-by-word from aligned corpora.
- Two flavors: Base (semantic speech tokens) and Expressive (adds pitch + style tokens) — an ancestor of today's natively speaking assistants.
Try it yourself
Code and checkpoint instructions (GitHub) ↗ Listen to generation samples ↗ Read the paper ↗