metaai·lightalo unofficial · independent
Universe / Safety & Trust / Purple Llama
Safety & Trust · 2023

Purple Llama

Purple Llama (Llama Guard, LlamaFirewall, Prompt Guard, CyberSecEval)

Open-source safety tooling for open models

open source Llama Guard 4: 12B; Prompt Guard 2: 22M / 86M params
★ 4.4k PurpleLlama❝ 1.3k citations⤓ 64k downloads/moread 2026-09-03

Latest: Llama Guard 4 12B + LlamaFirewall + Prompt Guard 2 (Apr 2025)

Meta's open trust-and-safety line, launched December 2023. Llama Guard classifies unsafe prompts/responses; Code Shield and CyberSecEval target insecure code and cyber risk. LlamaCon 2025 added Llama Guard 4 (a single 12B natively multimodal safeguard), LlamaFirewall (guardrail orchestration against prompt injection and risky tool use), and Prompt Guard 2.

Why it matters

Meta's open trust-and-safety toolkit: Llama Guard classifies unsafe prompts and responses, Prompt Guard catches jailbreaks and injections, CyberSecEval measures cyber risk. Llama Guard 4 is a single 12B natively multimodal safeguard pruned from Llama 4 Scout, making open guardrails a standard layer in Llama and third-party deployments.

Facts

Try it yourself

Lineage

Descends fromLlama 2
Led toCyberSecEvalLlamaFirewall

See the whole family tree →

Sources

More in Safety & Trust

GAIA Benchmark

Read the Safety & Trust story on the sky →

✦ Open on the map Explore Safety & Trust Quiz me