Train-before-Test Harmonizes Language Model Rankings

ICLROral2026

Authors
Guanhua Zhang, Ricardo Dominguez-Olmedo, Moritz Hardt
Affiliation
Max Planck Institute for Intelligent Systems, Max-Planck Institute
Venue
ICLR 2026
Track
Oral

TL;DR

Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model selection, clouds model comparisons, and adds confusion to a growing ecosystem of competing models.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model benchmark

← All ICLR 2026 Oral papers · Browse the whole archive