415.tech
AI & tech, from the frontlines of Silicon Valley
Willison's pelican benchmark no longer ranks frontier models

Willison's pelican benchmark no longer ranks frontier models

Simon Willison ran his 21-month-old pelican-on-a-bicycle SVG test on Moonshot's new Kimi K3 — a 2.8T-parameter model priced at $3/$15 per million tokens, matching Claude Sonnet — and found the benchmark has decoupled from actual model quality. GLM-5.2's pelican now outdraws Claude Fable 5 and GPT-5.6 despite GLM not being frontier-class, so the test survives as a cheap within-family check on cost and SVG geometry, not a cross-model ranking. Its blind spot is what now decides frontier rank: reliable agentic tool-calling as conversations grow long.

Source: simonwillison.net

Post on XEmail

The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.

Simon Willison

Why this matters

  • → SVG drawing test no longer correlates with frontier model ranking
  • → Agentic tool-calling and long-context reliability now define leader tier
  • → Benchmark survives as cheap within-family cost and geometry check
Benchmark obsolescence