
GPT-6 Astra clears ARC-AGI-3 levels in fewer moves than the median human, pulling Chollet's AGI forecast forward
Epoch AI ranks GPT-6 Astra first at 169 points across 50-plus benchmarks while Artificial Analysis scores it 61 — level with GPT-5.6 Sol and behind Claude Fable 5.1 at 66 — so the aggregate verdict is genuinely split. The uncontested result is ARC-AGI-3, where Astra reached 62.7% at roughly $26,000 in test cost against 7.8% for Sol and 30.2% for Claude Opus 5, and solved 96% of levels in fewer moves than the median human who cleared them. Francois Chollet called the progress twice as fast as expected and moved his 2030 AGI forecast sooner, shifting the frontier contest from raw benchmark scores to sample efficiency — how little experience a model needs before it masters an unfamiliar environment.
Source: the-decoder.com ↗
Astra cleared 96 percent of levels in fewer moves than the median human who solved them, on average with a little over half.
Why this matters
- → Sample efficiency now favors AI over humans—faster learning from fewer trials.
- → Chollet's AGI timeline compressed; frontier shifted from raw scores to environmental adaptation.
- → Model invents symbolic notation on-the-fly, internalizing reasoning that once required external harnesses.