415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI finds about a third of SWE-Bench Pro's coding tasks are broken

OpenAI finds about a third of SWE-Bench Pro's coding tasks are broken

OpenAI audited SWE-Bench Pro and found it unreliable — five human engineers flagged 249 of the 731 public tasks (34%) as broken, citing overly strict or misdirected hidden tests, including one that required two leading spaces where the prompt showed one. It is a concrete warning for anyone benchmarking coding agents: OpenAI walked back its own recommendation to adopt the benchmark, and frontier accuracy leaping from 23.3% to 80.3% in eight months points to contamination as much as real progress.

Source: the-decoder.com

Post on XEmail