P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling
ICLROral2026
TL;DR
The first personalized generative reward model with test-time user-based scaling for preference alignment…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
alignment