Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
ICLROral2026
TL;DR
Gaia2 evaluates LLM agents in asynchronous, dynamic environments with action-level verification, revealing fundamental trade-offs between reasoning, speed, and robustness.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
robustness benchmark reasoning agent llm