AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
ICLROral2026
TL;DR
We present principles and tooling for rigorous AI agent benchmarking, instantiated in AstaBench—the first holistic measure of agentic ability for scientific research—plus experiments showing AI remains far from solving research assistance.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
benchmark agent