
OpenAI revised GPT-6 Astra's benchmark numbers after launch, cutting rival scores
Fortune traced OpenAI editing published benchmarks in the hours after the GPT-6 Astra blog post went live: Astra's hallucination rate went 4.2% to 2% and back to 4.2%, its ARC-AGI-3 score rose from 98.6% in the embargoed draft to 99.99%, and Anthropic's Fable 5.1 FrontierMath Tier 4 score fell from 87.8% to 78% before landing at 83%. OpenAI attributes the swings to evaluation noise of a few percentage points across checkpoints, scaffolds, and runs — an explanation that also concedes vendor-published benchmark tables are not a stable basis for comparing frontier models.
Source: fortune.com ↗
Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting.
OpenAI spokesperson
Why this matters
- → Benchmark manipulation threatens fair model comparison in a high-stakes industry race
- → OpenAI's post-launch edits undermine credibility of vendor-published performance claims
- → Lack of transparency on evaluation conditions makes it impossible to trust reported numbers
Benchmarks under fire