Token-Importance Guided Direct Preference Optimization

ICLROral2026

Authors
Ning Yang, Hai Lin, Yibo Liu, Baoliang Tian, Guoqing Liu, Haijun Zhang
Affiliation
Institute of automation, Chinese academy of science, Chinese Academy of Sciences
Venue
ICLR 2026
Track
Oral

TL;DR

We proposes Token-Importance Guided Direct Preference Optimization (TI-DPO) to better align LLMs with human preferences by using a hybrid weighting mechanism to identify key tokens and a triplet loss to guide the optimization process.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

preference optimization optimization llm

← All ICLR 2026 Oral papers · Browse the whole archive