How Reliable is Language Model Micro-Benchmarking?
ICLROral2026
TL;DR
Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-benchmarks, however, rank models as consistently as the full benchmarks they replace?
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model benchmark