A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders

NeurIPSOral2025

Authors
David Chanin, James Wilken-Smith, Tomáš Dulka, Hardik Bhatnagar, Satvik Golechha, Joseph Isaac Bloom
Affiliation
MATS
Venue
NeurIPS 2025
Track
Oral

TL;DR

Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number of features in the SAE, hierarchical features tend to split into finer features (“math” may split into “algebra”, “geometry”, etc.), a phenomenon referred to as feature sp…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model llm

← All NeurIPS 2025 Oral papers · Browse the whole archive