
Anthropic raises its misalignment catastrophe estimate from 'very low' to 'low'
Anthropic's second company-wide Risk Report, covering February 24 to July 15, raises its estimate of catastrophic harm from misalignment in high-stakes settings from 'very low' to 'low' — an uncertainty adjustment after recent cybersecurity incident disclosures, not a new finding. Anthropic also discloses Model 2, an unreleased internal model somewhat more capable than Mythos 5, with no current plans for external release. A frontier lab is now publishing both a raised catastrophic-risk number and the existence of a model it is holding back.
Source: anthropic.com ↗