Hallucination Begins Where Saliency Drops
ICLROral2026
TL;DR
Recent studies have investigated attention dynamics in large vision language models (LVLMs), yet existing methods remain limited in reliably distinguishing hallucinated from correct outputs — primarily because they rely solely on forward-pass attention, ignoring gradient-based signals that reveal how token influence propagates through the model....
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model hallucination attention