415.tech
AI & tech, from the frontlines of Silicon Valley
Terminal-Bench-Science launches with 70 expert tasks; Claude Opus 5 leads at 30%

Terminal-Bench-Science launches with 70 expert tasks; Claude Opus 5 leads at 30%

Stanford researchers and the Terminal-Bench team released Terminal-Bench-Science 0.1, a benchmark of 70 expert-curated research workflows across life, physical, Earth, mathematical, and engineering sciences, distilled from 920 proposals by 376 contributors in 22 countries. Claude Opus 5 tops the leaderboard at 30.0% resolution, ahead of GPT-5.6 Sol at 22.4% and Claude Fable 5 at 21.4% — every model scores more than 10 points below its Terminal-Bench 3.0 mark, since tasks were deliberately calibrated to break frontier systems. The 70% failure rate is the honest read on agents as research assistants: real scientific workflows remain mostly unsolved, and the benchmark is continuous, so the bar moves with each release.

Source: terminal-bench-science.ai

Post on XEmail

Scientists, not model developers or data vendors, set the bar for scientific capability in AI.

Terminal-Bench-Science team

Why this matters

  • → Scientific benchmarks now measure real workflows, not textbook questions.
  • → 70% failure rate shows AI agents remain far from practical research assistance.
  • → Continuous benchmark creates feedback loop between scientific needs and frontier AI.
Science sets the bar
Also in this edition