ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs

NeurIPSSpotlight2025

Authors
Landon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas, Bin Yu, Kannan Ramchandran
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Large Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model interpretability language model efficient llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive