SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety

ICLROral2026

Authors
Geon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim, Honglak Lee, Kyunghoon Bae, Youngsoo Jang, Moontae Lee
Affiliation
LG AI Research
Venue
ICLR 2026
Track
Oral

TL;DR

This work introduces a simple yet principled approach for directly optimizing the safety alignment objective during policy learning…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

preference optimization optimization alignment safety

← All ICLR 2026 Oral papers · Browse the whole archive