Llama Stack
Write the app once, run it local, cloud or on-prem
Latest: v0.7.2 (May 28, 2026)
Open-source framework standardizing the APIs around Llama-based applications — inference, RAG, agents, tools, safety, evals, telemetry — with swappable providers so apps move between local, cloud, and on-prem unchanged. Launched September 2024 alongside Llama 3.2; now developed in its own llamastack GitHub org with partners like NVIDIA, IBM, Red Hat, and Dell.
Why it matters
Llama Stack tried to standardize the whole app layer around Llama — inference, RAG, agents, safety, evals — with swappable local and cloud providers, and drew NVIDIA, IBM, Red Hat and Dell as partners. In 2026 the project outgrew its name: it is now OGX, a model-agnostic, OpenAI-compatible agentic API server outside the Meta org.
Facts
- The 0.5 release (Feb 2026) added OpenAI API conformance and an MCP-server Connectors API — Meta's open stack absorbing industry standards.
- Now community-governed at llamastack/llama-stack after moving out of the meta-llama org.
Try it yourself
pip install llama-stack ↗ The project today: OGX on GitHub ↗ Why Llama Stack became OGX ↗
Lineage
Sources
GitHub · llamastack/llama-stack · llama stack ↗GitHub · llamastack/llama-stack · releases ↗pypi.org ↗