415.tech
AI & tech, from the frontlines of Silicon Valley
OpenAI revised GPT-6 Astra's benchmark numbers after launch, cutting rival scores

OpenAI revised GPT-6 Astra's benchmark numbers after launch, cutting rival scores

Fortune traced OpenAI editing published benchmarks in the hours after the GPT-6 Astra blog post went live: Astra's hallucination rate went 4.2% to 2% and back to 4.2%, its ARC-AGI-3 score rose from 98.6% in the embargoed draft to 99.99%, and Anthropic's Fable 5.1 FrontierMath Tier 4 score fell from 87.8% to 78% before landing at 83%. OpenAI attributes the swings to evaluation noise of a few percentage points across checkpoints, scaffolds, and runs — an explanation that also concedes vendor-published benchmark tables are not a stable basis for comparing frontier models.

Source: fortune.com

Post on XEmail

Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting.

OpenAI spokesperson

Why this matters

  • → Benchmark manipulation threatens fair model comparison in a high-stakes industry race
  • → OpenAI's post-launch edits undermine credibility of vendor-published performance claims
  • → Lack of transparency on evaluation conditions makes it impossible to trust reported numbers
Benchmarks under fire
Also in this edition