415.tech
AI & tech, from the frontlines of Silicon Valley
Anthropic proposes a CJS-0 to CJS-4 scale for rating AI cyber jailbreaks

Anthropic proposes a CJS-0 to CJS-4 scale for rating AI cyber jailbreaks

Anthropic proposed the Cyber Jailbreak Severity (CJS) scale — five levels from CJS-0 (informational) to CJS-4 (critical) — developed with its Project Glasswing partners to rate a jailbreak by capability gain, task breadth, weaponization ease, and discoverability. It is an early draft, not a standard, meant to give AI developers and governments shared language for jailbreak risk. It arrives with Fable 5's dual-use cyber classifiers and a new HackerOne bounty for researchers to submit jailbreaks.

Source: anthropic.com

Post on XEmail

Such a framework would allow AI developers to speak to governments (and vice versa) in consistent terms about the risks posed by each jailbreak.

Anthropic

Why this matters

  • → Establishes shared language for rating AI jailbreak risks across industry and government
  • → Enables defenders to use AI for cybersecurity while preventing attacker misuse
  • → Addresses dual-use challenge: same capabilities help both defenders and attackers
Jailbreak severity scale