415.tech
AI & tech, from the frontlines of Silicon Valley
Anthropic's automated researchers beat human proposals on 10 alignment benchmarks at $4/hour

Anthropic's automated researchers beat human proposals on 10 alignment benchmarks at $4/hour

Anthropic fellow Chen Yueh-Han's paper reports automated researchers improved a target model on all 10 alignment benchmarks tested without degrading overall performance, each iteration searching the literature, proposing a method, and training for 30 minutes. The best automated method beat what experienced humans propose within six hours, at roughly $4 per hour of API inference against the $150 per hour Anthropic pays human researchers — the first concrete cost curve for recursive self-improvement in alignment work. The paper concedes the result only holds insofar as the benchmarks reflect actual alignment goals, which leaves benchmark design as the remaining human job.

Source: techcrunch.com

Post on XEmail

The best AAR method beats what experienced humans propose, on average within six hours.

Anthropic paper, 'Automated Researchers Can Reliably Mitigate Alignment Failures'

Why this matters

  • → AI systems now beat human alignment researchers at 1/37th the cost.
  • → Opens path to recursive self-improvement in safety work.
  • → Benchmark design becomes the bottleneck, not research labor.
Researchers automate research
Also in this edition