Muse Glimmer
The open-source plot twist: a 30B agent that runs on your laptop
Latest: Muse Glimmer 30B (Aug 10, 2026)
A 30B-parameter open-weight multimodal agent model released August 10, 2026 under Apache 2.0 — Meta's return to open weights after Muse Spark went closed. Distilled from Muse Spark, it runs multi-step agent workflows (tool calls, coding, files, screenshots) entirely offline on a single consumer GPU or Mac, in 100+ languages.
Why it matters
Glimmer is Meta's answer to the charge that it abandoned open source: a 30B Apache-2.0 agent model distilled from Muse Spark that runs offline on one consumer GPU and shipped day-one on Hugging Face, Ollama, llama.cpp and MLX. It re-established Meta as an open-weight player after Llama 4.
Facts
- Quantizes to under 20 GB.
- Speculative decoding (DFlash drafter) gives 1.5–3.1x faster generation.
- Optimized for MacBook M4-Max/M5-Max and RTX-5090 with AMD, Arm, Dell, Intel, NVIDIA as partners.
- Ships on Hugging Face, Ollama, LM Studio, llama.cpp, ExecuTorch, MLX, vLLM, SGLang.
- Launched alongside a public recommitment from Zuckerberg to open source, with an open-weight version of Muse Spark 1.2 promised next.
Try it yourself
Model on Hugging Face ↗ Run it locally with Ollama ↗ Use it via OpenRouter ↗
Lineage
Sources
Hugging Face · Muse Glimmer 30B ↗Hugging Face · muse glimmer ↗Meta Research ↗Meta for Developers ↗CNBC ↗TechCrunch ↗Fortune ↗