EigenBench: A Comparative Behavioral Measure of Value Alignment
ICLROral2026
TL;DR
Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models’ values.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model alignment benchmark