VPO: Reasoning Preferences Optimization Based on $\mathcal{V}$-Usable Information
NeurIPSSpotlight2025
TL;DR
Direct Preference Optimization (DPO) is a widely used preference optimization algorithm in large language model (LLM) alignment, which reparameterizes the reward function in reinforcement learning wit…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
preference optimization reinforcement learning large language model language model optimization alignment reasoning llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive