
OpenAI finds about a third of SWE-Bench Pro's coding tasks are broken
OpenAI audited SWE-Bench Pro and found it unreliable — five human engineers flagged 249 of the 731 public tasks (34%) as broken, citing overly strict or misdirected hidden tests, including one that required two leading spaces where the prompt showed one. It is a concrete warning for anyone benchmarking coding agents: OpenAI walked back its own recommendation to adopt the benchmark, and frontier accuracy leaping from 23.3% to 80.3% in eight months points to contamination as much as real progress.
Source: the-decoder.com ↗