The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization

NeurIPSSpotlight2025

Authors
Yihua Zhang, Changsheng Wang, Yiwei Chen, Chongyu Fan, Jinghan Jia, Sijia Liu
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Input saliency aims to quantify the influence of input tokens on the output of large language models (LLMs), which has been widely used for prompt engineering, model interpretability, and behavior att…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model interpretability language model optimization attention llm rag

← All NeurIPS 2025 Spotlight papers · Browse the whole archive