What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

ICLROral2026

Authors
Rajiv Movva, Smitha Milli, Sewon Min, Emma Pierson
Affiliation
University of California, Berkeley
Venue
ICLR 2026
Track
Oral

TL;DR

We present WIMHF, a method to describe the preferences encoded by human feedback; produce insights from seven widely-used datasets; and show that the method enables new approaches to data curation and personalization.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

dataset

← All ICLR 2026 Oral papers · Browse the whole archive