Bits Leaked per Query: Information-Theoretic Bounds for Adversarial Attacks on LLMs

NeurIPSSpotlight2025

Authors
Masahiro Kaneko, Timothy Baldwin
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Adversarial attacks by malicious users that threaten the safety of large language models (LLMs) can be viewed as attempts to infer a target property $T$ that is unknown when an instruction is issued,…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model adversarial safety llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive