
Terminal-Bench-Science launches with 70 expert tasks; Claude Opus 5 leads at 30%
Stanford researchers and the Terminal-Bench team released Terminal-Bench-Science 0.1, a benchmark of 70 expert-curated research workflows across life, physical, Earth, mathematical, and engineering sciences, distilled from 920 proposals by 376 contributors in 22 countries. Claude Opus 5 tops the leaderboard at 30.0% resolution, ahead of GPT-5.6 Sol at 22.4% and Claude Fable 5 at 21.4% — every model scores more than 10 points below its Terminal-Bench 3.0 mark, since tasks were deliberately calibrated to break frontier systems. The 70% failure rate is the honest read on agents as research assistants: real scientific workflows remain mostly unsolved, and the benchmark is continuous, so the bar moves with each release.
Source: terminal-bench-science.ai ↗
Scientists, not model developers or data vendors, set the bar for scientific capability in AI.
Why this matters
- → Scientific benchmarks now measure real workflows, not textbook questions.
- → 70% failure rate shows AI agents remain far from practical research assistance.
- → Continuous benchmark creates feedback loop between scientific needs and frontier AI.