Invisible Safety Threat: Malicious Finetuning for LLM via Steganography
ICLROral2026
TL;DR
We highlight an insidious safety threat: a compromised LLM can maintain a facade of proper safety alignment while covertly generating harmful content through steganography.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
alignment safety graph gan llm