415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI limits Astra's advanced cyber features to a small alpha group

OpenAI limits Astra's advanced cyber features to a small alpha group

Astra is the first OpenAI model to cross the Critical cybersecurity threshold in its Preparedness Framework, beating GPT-5.6 Sol on ExploitBench's 20 high-severity vulnerabilities and chaining two zero-days it discovered itself. Full cyber access goes first to a small alpha group including the US government and critical-infrastructure defenders, widening later via the Daybreak Blue program — autonomous exploit discovery is now a gated commercial product, not a lab demo. Astra refused 91.5% of inappropriate cyber requests against Sol's 59%, which still leaves 8.5% complied with.

Source: fortune.com

Post on XEmail

Astra is the first model it plans to release that meets its "critical cybersecurity capability threshold" under its Preparedness Framework, an internal policy that governs the safety precautions the company will put in place depending on the risks a model presents.

OpenAI

Why this matters

  • → Autonomous AI exploit discovery moves from lab demo to gated commercial product with real government deploymen
  • → 8.5% compliance rate on harmful requests reveals fundamental tradeoff: safety refusals block legitimate defens
Defense or dominance
Also in this edition