EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits

ICLROral2026

Authors
Wayne Chi, Valerie Chen, Ryan Shar, Aditya Mittal, Jenny Liang, Wei-Lin Chiang, Anastasios Nikolas Angelopoulos, Ion Stoica, Graham Neubig, Ameet Talwalkar, Chris Donahue
Affiliation
Carnegie Mellon University
Venue
ICLR 2026
Track
Oral

TL;DR

We propose a new benchmark for evaluating an LLM's ability to perform code edits. Our data is gathered from in-the-wild code edits, leading to more realistic problems.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

benchmark llm

← All ICLR 2026 Oral papers · Browse the whole archive