EigenBench: A Comparative Behavioral Measure of Value Alignment

ICLROral2026

Authors
Jonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine Xinze Li, Lionel Levine
Affiliation
Cornell University
Venue
ICLR 2026
Track
Oral

TL;DR

Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models’ values.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model alignment benchmark

← All ICLR 2026 Oral papers · Browse the whole archive