The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology

ICLROral2026

Authors
Aideen Fay, Inés García-Redondo, Qiquan Wang, Haim Dubossarsky, Anthea Monod
Affiliation
Imperial College London
Venue
ICLR 2026
Track
Oral

TL;DR

We use persistent homology to interpret how adversarial inputs reshape LLM representation spaces, resulting in a robust signature that provides multiscale, geometry-aware insights complementary to standard interpretability methods.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

interpretability adversarial llm

← All ICLR 2026 Oral papers · Browse the whole archive