Semi-Supervised Preference Optimization with Limited Feedback

ICLROral2026

Authors
Seonggyun Lee, Sungjun Lim, Seojin Park, Soeun Cheon, Kyungwoo Song
Affiliation
Yonsei University
Venue
ICLR 2026
Track
Oral

TL;DR

The field of preference optimization has made outstanding contributions to the alignment of language models with human preferences. Despite these advancements, recent methods still rely heavily on substantial paired (labeled) feedback data, leading to substantial resource expenditures.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

preference optimization language model optimization alignment

← All ICLR 2026 Oral papers · Browse the whole archive