
Anthropic proposes a CJS-0 to CJS-4 scale for rating AI cyber jailbreaks
Anthropic proposed the Cyber Jailbreak Severity (CJS) scale — five levels from CJS-0 (informational) to CJS-4 (critical) — developed with its Project Glasswing partners to rate a jailbreak by capability gain, task breadth, weaponization ease, and discoverability. It is an early draft, not a standard, meant to give AI developers and governments shared language for jailbreak risk. It arrives with Fable 5's dual-use cyber classifiers and a new HackerOne bounty for researchers to submit jailbreaks.
Source: anthropic.com ↗
Such a framework would allow AI developers to speak to governments (and vice versa) in consistent terms about the risks posed by each jailbreak.
Anthropic
Why this matters
- → Establishes shared language for rating AI jailbreak risks across industry and government
- → Enables defenders to use AI for cybersecurity while preventing attacker misuse
- → Addresses dual-use challenge: same capabilities help both defenders and attackers
Jailbreak severity scale