Train-before-Test Harmonizes Language Model Rankings
ICLROral2026
TL;DR
Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model selection, clouds model comparisons, and adds confusion to a growing ecosystem of competing models.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model benchmark