Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

ICLROral2026

Authors
Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev
Affiliation
ETH Zurich
Venue
ICLR 2026
Track
Oral

TL;DR

We show that adversaries can implant hidden adversarial behaviors in LLM that are inadvertently triggered by users finetuning the model.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

adversarial llm

← All ICLR 2026 Oral papers · Browse the whole archive