The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
NeurIPSSpotlight2025
TL;DR
The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
interpretability causal
← All NeurIPS 2025 Spotlight papers · Browse the whole archive