
Anthropic's automated researchers beat human proposals on 10 alignment benchmarks at $4/hour
Anthropic fellow Chen Yueh-Han's paper reports automated researchers improved a target model on all 10 alignment benchmarks tested without degrading overall performance, each iteration searching the literature, proposing a method, and training for 30 minutes. The best automated method beat what experienced humans propose within six hours, at roughly $4 per hour of API inference against the $150 per hour Anthropic pays human researchers — the first concrete cost curve for recursive self-improvement in alignment work. The paper concedes the result only holds insofar as the benchmarks reflect actual alignment goals, which leaves benchmark design as the remaining human job.
Source: techcrunch.com ↗
The best AAR method beats what experienced humans propose, on average within six hours.
Why this matters
- → AI systems now beat human alignment researchers at 1/37th the cost.
- → Opens path to recursive self-improvement in safety work.
- → Benchmark design becomes the bottleneck, not research labor.