Among Us: A Sandbox for Measuring and Detecting Agentic Deception

NeurIPSSpotlight2025

Authors
Satvik Golechha, Adrià Garriga-Alonso
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Prior studies on deception in language-based AI agents typically assess whether the agent produces a false statement about a topic, or makes a binary choice prompted by a goal, rather than allowing op…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

agent

← All NeurIPS 2025 Spotlight papers · Browse the whole archive