SWINGARENA: Adversarial Programming Arena for Long-context GitHub Issue Solving
ICLROral2026
TL;DR
We present \textsc{SwingArena}, a adversarial evaluation framework for Large Language Models (LLMs) that closely mirrors real-world software development workflows. Unlike traditional static benchmarks, \textsc{SwingArena} models the collaborative process of software iteration by pairing LLMs as \textit{submitters}, who generate patches, and \tex...
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model adversarial evaluation benchmark llm