CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale

ICLROral2026

Authors
Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, Dawn Song
Venue
ICLR 2026
Track
Oral

TL;DR

AI agents have significant potential to reshape cybersecurity, making a thorough assessment of their capabilities critical. However, existing evaluations fall short, because they are based on small-scale benchmarks and only measure static outcomes, failing to capture the full, dynamic range of real-world security challenges.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

evaluation benchmark agent

← All ICLR 2026 Oral papers · Browse the whole archive