Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
ICLROral2026
TL;DR
We propose that using contextual information to train SAEs will improve their representation of semantic and high-level features.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
interpretability rag