415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI tops a five-lab study on rogue-model containment plans with 3 out of 5

OpenAI tops a five-lab study on rogue-model containment plans with 3 out of 5

Guidelight AI Standards graded Anthropic, Google, Meta, OpenAI, and xAI on published plans for containing a model that escapes human control — OpenAI led at 3 out of 5, Anthropic and Meta scored lowest. The grades measure public disclosure rather than internal safeguards, but the gap is now a compliance matter: California's SB 53 already requires frontier developers to publish incident-response frameworks, and New York's RAISE Act takes effect in January.

Source: techcrunch.com

Post on XEmail

I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense.

Steven Adler, Guidelight AI Standards chief scientist

Why this matters

  • → Regulators now mandate public containment plans—California's SB 53 already active, New York's RAISE Act starts
  • → Most AI labs scored poorly on emergency shutdown capabilities as agentic models gain autonomous access to comp
  • → Gap between public safety claims and actual incident-response readiness is now a compliance and liability issu
Kill switch compliance gap
Also in this edition