Among Us: A Sandbox for Measuring and Detecting Agentic Deception
NeurIPSSpotlight2025
TL;DR
Prior studies on deception in language-based AI agents typically assess whether the agent produces a false statement about a topic, or makes a binary choice prompted by a goal, rather than allowing op…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
agent
← All NeurIPS 2025 Spotlight papers · Browse the whole archive