Llama API
Meta's hosted Llama, folded into the Model API in 2026
Latest: Public preview (Apr 2025) — folded into the Meta Model API (Jul 2026)
Meta's first-party hosted inference service, previewed at LlamaCon (April 29, 2025) with OpenAI-SDK compatibility, one-click keys, and fine-tuning tools, plus fast-inference partnerships with Cerebras and Groq. It never left preview: in July 2026 it was folded into the newer Muse-era Meta Model API, with developers steered to third-party hosts for Llama.
Why it matters
The Llama API was Meta's first attempt at hosted inference — OpenAI-SDK compatible, one-click keys, Cerebras and Groq fast paths — announced as Llama passed one billion downloads. It never left preview; its 15-month arc from LlamaCon to absorption into the Muse-era Model API traces Meta's pivot from open Llama to paid Muse.
Facts
- Meta pledged 'we do not use your prompts or model responses to train our AI models.' Its 15-month life (preview to sunset) neatly brackets Meta's pivot from Llama to Muse.
Try it yourself
Read the LlamaCon announcement ↗ Its successor: Meta Model API docs ↗ Llama 4 hosted on OpenRouter today ↗
Lineage
Sources
Meta AI blog ↗Meta for Developers ↗