How Reliable is Language Model Micro-Benchmarking?

ICLROral2026

Authors
Gregory Yauney, Shahzaib Saqib Warraich, Swabha Swayamdipta
Affiliation
Amazon
Venue
ICLR 2026
Track
Oral

TL;DR

Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-benchmarks, however, rank models as consistently as the full benchmarks they replace?

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model benchmark

← All ICLR 2026 Oral papers · Browse the whole archive