Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

ICMLOral2026

Authors
Zhibin Duan, Guowei Rong, Zhuo Li, Bo Chen, Mingyuan Zhou, Dandan Guo
Venue
ICML 2026
Track
Oral

TL;DR

No summary has been collected for this paper yet. Read the abstract at the authoritative source below.

Read the paper

Topics

bayesian rlhf

← All ICML 2026 Oral papers · Browse the whole archive