Invisible Safety Threat: Malicious Finetuning for LLM via Steganography

ICLROral2026

Authors
Guangnian Wan, Xinyin Ma, Gongfan Fang, Xinchao Wang
Affiliation
National University of Singapore
Venue
ICLR 2026
Track
Oral

TL;DR

We highlight an insidious safety threat: a compromised LLM can maintain a facade of proper safety alignment while covertly generating harmful content through steganography.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

alignment safety graph gan llm

← All ICLR 2026 Oral papers · Browse the whole archive