415.tech
AI & tech, from the frontlines of Silicon Valley
Artificial Analysis doubles private held-out data to 40% of its Intelligence Index in v4.2

Artificial Analysis doubles private held-out data to 40% of its Intelligence Index in v4.2

Artificial Analysis shipped Intelligence Index v4.2, doubling private held-out test sets to 40% of the weighting, adding AA-Briefcase for agentic knowledge work and Surge AI's GDP.pdf across 4,592 PDF pages, and dropping the saturated GPQA Diamond. The anti-gaming shift is the real change — leaderboard position now rests on data labs cannot train against. Claude Fable 5.1 still leads overall, while GPT-6 Astra takes second and tops GDP.pdf at 33.2% to Fable 5.1's 26.2%.

Source: artificialanalysis.ai

Post on XEmail

40% of our Index weighting is now private, held-out test sets - double the figure from v4.1.

Artificial Analysis

Why this matters

  • → 40% private test sets block model lab gaming via public benchmarks
  • → Agentic knowledge work now measured as realistic multi-week project performance
  • → Frontier leaders (Fable 5.1, GPT-6 Astra) diverge sharply on long-context document reasoning
Gaming-proof benchmarks
Also in this edition