The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?

NeurIPSSpotlight2025

Authors
Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

interpretability causal

← All NeurIPS 2025 Spotlight papers · Browse the whole archive