The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
ICLROral2026
TL;DR
We use persistent homology to interpret how adversarial inputs reshape LLM representation spaces, resulting in a robust signature that provides multiscale, geometry-aware insights complementary to standard interpretability methods.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
interpretability adversarial llm