SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
ICLROral2026
TL;DR
This work introduces a simple yet principled approach for directly optimizing the safety alignment objective during policy learning…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
preference optimization optimization alignment safety