Reliable Weak-to-Strong Monitoring of LLM Agents

ICLROral2026

Authors
Neil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich, Paula Rodriguez, Christina Q Knight, Zifan Wang
Affiliation
School of Computer Science, Carnegie Mellon University
Venue
ICLR 2026
Track
Oral

TL;DR

This paper introduces a monitor red teaming workflow to stress test systems for detecting covert misbehavior in LLM agents, finding that a well-designed monitor scaffold enables weaker models to oversee strong aware attackers.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

agent llm

← All ICLR 2026 Oral papers · Browse the whole archive