Token-Importance Guided Direct Preference Optimization
ICLROral2026
TL;DR
We proposes Token-Importance Guided Direct Preference Optimization (TI-DPO) to better align LLMs with human preferences by using a hybrid weighting mechanism to identify key tokens and a triplet loss to guide the optimization process.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
preference optimization optimization llm