Multiplayer Nash Preference Optimization
ICLROral2026
TL;DR
Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley–Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
preference optimization reinforcement learning large language model language model optimization rlhf