metaai·lightalo unofficial · independent
Universe / Safety & Trust / CyberSecEval
Safety & Trust · 2023

CyberSecEval

Can your model break code — and can it patch it?

open source
★ 4.4k PurpleLlamaread 2026-09-03

Latest: CyberSecEval 4 (April 2025)

The most comprehensive open benchmark suite for measuring LLM cybersecurity risks and capabilities, spanning insecure code generation, cyberattack helpfulness, prompt-injection resistance and vulnerability exploitation. CyberSecEval 4 (April 2025) added defensive evaluations: AutoPatchBench (automatic vulnerability patching) and CyberSOCEval, built with CrowdStrike for SOC-style malware and threat-intel analysis.

Why it matters

The most comprehensive open benchmark for LLM cybersecurity risk: insecure code generation, cyberattack helpfulness, prompt-injection resistance and vulnerability exploitation. CyberSecEval 4 added defensive evaluations, AutoPatchBench and CyberSOCEval with CrowdStrike, and it has become a standard reference for cyber-safety evaluation of frontier models.

Facts

Try it yourself

Lineage

Descends fromPurple Llama

See the whole family tree →

Sources

More in Safety & Trust

GAIA BenchmarkLlamaFirewall

Read the Safety & Trust story on the sky →

✦ Open on the map Explore Safety & Trust Quiz me