metaai·lightalo unofficial · independent
Universe / Safety & Trust / LlamaFirewall
Safety & Trust · 2025

LlamaFirewall

Prompt-injection success rate: 17.6% without, 1.7% with

open source

Latest: LlamaFirewall (2025), part of the PurpleLlama repository

An open-source, system-level guardrail framework for securing AI agents, released April-May 2025 and used in production at Meta. Its three layers — PromptGuard 2 (jailbreak detection), Agent Alignment Checks (a chain-of-thought auditor catching goal hijacking), and CodeShield (static analysis of generated code) — defend against prompt injection and insecure agent behavior.

Why it matters

A system-level guardrail framework for AI agents, used in production at Meta and released open source: PromptGuard 2 for jailbreak detection, Agent Alignment Checks that audit chain-of-thought for goal hijacking, and CodeShield static analysis of generated code. One of the first open, layered defenses built specifically for agent workflows.

Facts

Try it yourself

Lineage

Descends fromPurple Llama

See the whole family tree →

Sources

More in Safety & Trust

CyberSecEvalGAIA Benchmark

Read the Safety & Trust story on the sky →

✦ Open on the map Explore Safety & Trust Quiz me