Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
ICLROral2026
TL;DR
We show that adversaries can implant hidden adversarial behaviors in LLM that are inadvertently triggered by users finetuning the model.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
adversarial llm