Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments

ICLROral2026

Authors
Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Laurençon, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre Menard, Gerard Moreno-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Grégoire Mialon, Thomas Scialom
Affiliation
Facebook
Venue
ICLR 2026
Track
Oral

TL;DR

Gaia2 evaluates LLM agents in asynchronous, dynamic environments with action-level verification, revealing fundamental trade-offs between reasoning, speed, and robustness.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

robustness benchmark reasoning agent llm

← All ICLR 2026 Oral papers · Browse the whole archive