P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

ICLROral2026

Authors
Pinyi Zhang, Ting-En Lin, Yuchuan Wu, Jingyang Chen, Zongqi Wang, Hua Yang, Xu Ze, Fei Huang, Yongbin Li, Kai Zhang
Affiliation
East China Normal University
Venue
ICLR 2026
Track
Oral

TL;DR

The first personalized generative reward model with test-time user-based scaling for preference alignment…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

alignment

← All ICLR 2026 Oral papers · Browse the whole archive