ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
NeurIPSSpotlight2025
TL;DR
Large Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model interpretability language model efficient llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive