
Artificial Analysis doubles private held-out data to 40% of its Intelligence Index in v4.2
Artificial Analysis shipped Intelligence Index v4.2, doubling private held-out test sets to 40% of the weighting, adding AA-Briefcase for agentic knowledge work and Surge AI's GDP.pdf across 4,592 PDF pages, and dropping the saturated GPQA Diamond. The anti-gaming shift is the real change — leaderboard position now rests on data labs cannot train against. Claude Fable 5.1 still leads overall, while GPT-6 Astra takes second and tops GDP.pdf at 33.2% to Fable 5.1's 26.2%.
Source: artificialanalysis.ai ↗
40% of our Index weighting is now private, held-out test sets - double the figure from v4.1.
Artificial Analysis
Why this matters
- → 40% private test sets block model lab gaming via public benchmarks
- → Agentic knowledge work now measured as realistic multi-week project performance
- → Frontier leaders (Fable 5.1, GPT-6 Astra) diverge sharply on long-context document reasoning
Gaming-proof benchmarks