
OpenAI tops a five-lab study on rogue-model containment plans with 3 out of 5
Guidelight AI Standards graded Anthropic, Google, Meta, OpenAI, and xAI on published plans for containing a model that escapes human control — OpenAI led at 3 out of 5, Anthropic and Meta scored lowest. The grades measure public disclosure rather than internal safeguards, but the gap is now a compliance matter: California's SB 53 already requires frontier developers to publish incident-response frameworks, and New York's RAISE Act takes effect in January.
Source: techcrunch.com ↗
I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense.
Steven Adler, Guidelight AI Standards chief scientist
Why this matters
- → Regulators now mandate public containment plans—California's SB 53 already active, New York's RAISE Act starts
- → Most AI labs scored poorly on emergency shutdown capabilities as agentic models gain autonomous access to comp
- → Gap between public safety claims and actual incident-response readiness is now a compliance and liability issu
Kill switch compliance gap