Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts

ICLROral2026

Authors
Zhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng He
Affiliation
National University of Singapore
Venue
ICLR 2026
Track
Oral

TL;DR

We detected the widespread deception of LLM under benign prompts and found its tendency increases with task difficulty.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

llm

← All ICLR 2026 Oral papers · Browse the whole archive