How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability

ICLROral2026

Authors
Shawn Im, Changdae Oh, Zhen Fang, Sharon Li
Affiliation
Department of Computer Science, University of Wisconsin - Madison
Venue
ICLR 2026
Track
Oral

TL;DR

Semantic associations such as the link between "bird" and "flew" are foundational for language modeling as they enable models to go beyond memorization and instead generalize and generate coherent text. Understanding how these associations are learned and represented in language models is essential for connecting deep learning with linguistic th...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

interpretability language model transformer

← All ICLR 2026 Oral papers · Browse the whole archive