EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
ICLROral2026
TL;DR
We propose a new benchmark for evaluating an LLM's ability to perform code edits. Our data is gathered from in-the-wild code edits, leading to more realistic problems.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
benchmark llm