metaai·lightalo unofficial · independent
Universe / Language & LLMs / Llama Stack
Language & LLMs · 2024

Llama Stack

Write the app once, run it local, cloud or on-prem

open source

Latest: v0.7.2 (May 28, 2026)

Open-source framework standardizing the APIs around Llama-based applications — inference, RAG, agents, tools, safety, evals, telemetry — with swappable providers so apps move between local, cloud, and on-prem unchanged. Launched September 2024 alongside Llama 3.2; now developed in its own llamastack GitHub org with partners like NVIDIA, IBM, Red Hat, and Dell.

Why it matters

Llama Stack tried to standardize the whole app layer around Llama — inference, RAG, agents, safety, evals — with swappable local and cloud providers, and drew NVIDIA, IBM, Red Hat and Dell as partners. In 2026 the project outgrew its name: it is now OGX, a model-agnostic, OpenAI-compatible agentic API server outside the Meta org.

Facts

Try it yourself

Lineage

Descends fromLlama 3 / 3.1 / 3.2 / 3.3

See the whole family tree →

Sources

More in Language & LLMs

Large Concept ModelsByte Latent TransformerCoconutMulti-token predictionMemory Layers at ScaleLLaMA

Read the Language & LLMs story on the sky →

✦ Open on the map Explore Language & LLMs Quiz me