PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation Models

NeurIPSOral2025

Authors
Ruiqi Wang, Dezhong Zhao, Ziqin Yuan, Tianyu Shao, Guohua Chen, Dominic Kao, Sungeun Hong, Byung-Cheol Min
Affiliation
Purdue University
Venue
NeurIPS 2025
Track
Oral

TL;DR

Preference-based reinforcement learning (PbRL) has emerged as a promising paradigm for teaching robots complex behaviors without reward engineering. However, its effectiveness is often limited by two critical challenges: the reliance on extensive human input and the inherent difficulties in resolving query ambiguity and credit assignment during rew…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning multimodal

← All NeurIPS 2025 Oral papers · Browse the whole archive