Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

ICMLOral2026

Authors
Yanchen Yin, Dongqi Han, Linghui Li
Venue
ICML 2026
Track
Oral

TL;DR

No summary has been collected for this paper yet. Read the abstract at the authoritative source below.

Read the paper

Topics

large language model language model attention

← All ICML 2026 Oral papers · Browse the whole archive