Multiplayer Nash Preference Optimization

ICLROral2026

Authors
Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, Xiaomin Li, Bing Hu, Peng Xia, Jure Leskovec, Yejin Choi
Venue
ICLR 2026
Track
Oral

TL;DR

Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grounded in the Bradley–Terry assumption struggle to capture the nontransitivity and heterogeneity of real-world preferences.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

preference optimization reinforcement learning large language model language model optimization rlhf

← All ICLR 2026 Oral papers · Browse the whole archive