415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI's GPT-Red finds prompt-injection exploits at 84% vs 13% for human red-teamers

OpenAI's GPT-Red finds prompt-injection exploits at 84% vs 13% for human red-teamers

OpenAI trained GPT-Red, an internal self-play RL system, to auto-discover prompt-injection exploits, hitting 84% on held-out indirect-injection tests against GPT-5.1 versus roughly 13% for human red-teamers. It surfaced a new attack class, "fake chain of thought," that spoofs a model's visible reasoning and fooled GPT-5.1 over 95% of the time — then dropped below 10% after adversarial training on GPT-5.6 Sol. Keeping the tool internal-only turns red-teaming into a closed loop where automated attack generation directly feeds the next model's defense, a capability rivals can't buy off the shelf.

Source: openai.com

Post on XEmail