Llama 4
Llama 4 (Scout / Maverick / Behemoth)
A 10-million-token context window — and a turning point
Latest: Scout & Maverick (Apr 5, 2025); Behemoth never shipped — no Llama 4.x/5 followed
Meta's last open-weight frontier Llama generation, released April 5, 2025: natively multimodal mixture-of-experts models. Scout (17B active, 16 experts, 109B total) shipped with an unprecedented 10M-token context window; Maverick (17B active, 128 experts, 400B total) targeted flagship quality. Behemoth (~2T parameters, 288B active) was previewed but never released.
Why it matters
Llama 4 was Meta's last open frontier generation and the moment the strategy cracked: Scout's 10M-token context and native multimodality were genuine firsts for open weights, but the 2T-parameter Behemoth never shipped amid reported MoE-routing problems, and within a year Meta replaced the line with the closed Muse Spark.
Facts
- A context window big enough to read the entire Harry Potter series about six times over in one prompt.
- Scout's 10M-token context was the largest of any open-weight model at launch.
- Llama downloads passed 1 billion in March 2025 (~1.2B by LlamaCon that April).
- Behemoth was used internally as a teacher to co-distill Scout and Maverick but its weights never shipped — mid-training MoE-routing issues at 2T scale are the reported cause.
- The line was superseded by the closed Muse Spark in April 2026.
Try it yourself
Llama 4 Scout on Hugging Face ↗ Run it locally with Ollama ↗ Use Maverick via OpenRouter ↗
Lineage
Sources
GitHub · llama-models ↗Hugging Face ↗Meta AI blog ↗llama.com ↗Axios ↗