Purple Llama
Purple Llama (Llama Guard, LlamaFirewall, Prompt Guard, CyberSecEval)
Open-source safety tooling for open models
Latest: Llama Guard 4 12B + LlamaFirewall + Prompt Guard 2 (Apr 2025)
Meta's open trust-and-safety line, launched December 2023. Llama Guard classifies unsafe prompts/responses; Code Shield and CyberSecEval target insecure code and cyber risk. LlamaCon 2025 added Llama Guard 4 (a single 12B natively multimodal safeguard), LlamaFirewall (guardrail orchestration against prompt injection and risky tool use), and Prompt Guard 2.
Why it matters
Meta's open trust-and-safety toolkit: Llama Guard classifies unsafe prompts and responses, Prompt Guard catches jailbreaks and injections, CyberSecEval measures cyber risk. Llama Guard 4 is a single 12B natively multimodal safeguard pruned from Llama 4 Scout, making open guardrails a standard layer in Llama and third-party deployments.
Facts
- The name is a security pun: 'purple teaming' = red (attack) + blue (defend).
- Llama Guard 4 unified separate text and vision guard models into one 12B model.
- The PurpleLlama repo has ~4,400 GitHub stars; an AI Defenders Program gives partners early security tooling.
- Llama Guard 4 runs on a single GPU and handles up to five images per prompt; the Llama Guard lineage (v1→v4 in 17 months) became the default open moderation layer across the industry.
Try it yourself
Llama Guard 4 on Hugging Face ↗ Prompt Guard 2 on Hugging Face ↗ Code on GitHub ↗
Lineage
Sources
arXiv ↗GitHub · PurpleLlama ↗Hugging Face ↗Meta AI blog · ai defenders program llama pro ↗Meta AI blog · purple llama open trust safety ↗llama.com ↗