metaai·lightalo unofficial · independent
Universe / Speech & Sound / Spirit LM
Speech & Sound · 2024

Spirit LM

Text and speech in one token stream — with the emotion intact

open source 7B params
★ 929 spiritlm❝ 156 citationsread 2026-09-03

Latest: Spirit LM Base & Expressive 7B (weights Oct 18, 2024)

Meta's first open multimodal language model that freely mixes text and speech in one token stream (paper Feb 2024, weights Oct 2024). The Expressive variant adds pitch and style tokens so it can carry emotion — anger, surprise, excitement — across modalities, enabling speech that actually sounds like it means it. FAIR Noncommercial Research License.

Why it matters

Meta's first open model that interleaves spoken and written tokens in one stream, so it can continue a text prompt in speech or the reverse, and its Expressive variant carries pitch and style across modalities. It is a reference design for speech-native language models, released under a non-commercial research license.

Facts

Try it yourself

Sources

More in Speech & Sound

SeamlessM4T & the Seamless familyOmnilingual ASRMassively Multilingual SpeechAudioCraftSAM AudioLanguage Technology Partner Program + BOUQuET

Read the Speech & Sound story on the sky →

✦ Open on the map Explore Speech & Sound Quiz me