
Databricks' own coding benchmark ties GLM 5.2 with Opus 4.8 at $1.28 vs $1.94 per task
Databricks benchmarked coding agents on its own multi-million-line codebase and found Chinese open-source GLM 5.2 statistically tied Opus 4.8 on task quality — 87% each — at $1.28 per task against Opus's $1.94. Token price proved a poor cost guide: Sonnet 5 runs ~1.7x cheaper per token yet cost $2.09 per task, because it read more and burned 1.9x the tokens to finish. Any team with a backlog of merged PRs can build the same benchmark — tasks no model trained on, graded by the team's own tests.
Source: databricks.com ↗
GLM can be a daily driver model for a lot of our developers. It landed in the top capability tier, statistically tied with Opus 4.8 on quality, but costing $1.28/task against Opus's $1.94.
Databricks engineering team
Why this matters
- → Open-source GLM matches Claude Opus quality at 34% lower cost per task
- → Token price is poor predictor of actual task cost due to reasoning efficiency variance
- → Teams can build production-grade benchmarks from their own merged PRs without external data
Open-source ties premium AI