
Willison's pelican benchmark no longer ranks frontier models
Simon Willison ran his 21-month-old pelican-on-a-bicycle SVG test on Moonshot's new Kimi K3 — a 2.8T-parameter model priced at $3/$15 per million tokens, matching Claude Sonnet — and found the benchmark has decoupled from actual model quality. GLM-5.2's pelican now outdraws Claude Fable 5 and GPT-5.6 despite GLM not being frontier-class, so the test survives as a cheap within-family check on cost and SVG geometry, not a cross-model ranking. Its blind spot is what now decides frontier rank: reliable agentic tool-calling as conversations grow long.
Source: simonwillison.net ↗
The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.
Why this matters
- → SVG drawing test no longer correlates with frontier model ranking
- → Agentic tool-calling and long-context reliability now define leader tier
- → Benchmark survives as cheap within-family cost and geometry check