acceptodds
Under review as a conference paper at ICLR 2027

Training Agent Monitors with Minimal Counterfactual Pairs

Abstract

Agent monitors support safer deployment by detecting malicious actions within legitimate work, sometimes before task completion. Yet separately collected benign and malicious runs can differ throughout execution, making it difficult to construct matched contrasts for training and evaluation. We propose a framework for training agent monitors with controlled minimal counterfactual pairs. We construct each pair by producing a malicious twin from a real benign trajectory through bounded local edits. An adversarial editor inserts actions for a malicious side task, while mechanical checks and behavioral judges assess structural validity and coherence. The construction also supports feedback-guided collection, using monitor responses to target subsequent edits. The pairs support direct monitor training and controlled trajectory-level, paired, and prefix-level evaluation. On CUA-SHADE-Arena, an external benchmark of computer-use agent trajectories, direct training on our pairs raises Qwen3.5-35B-A3B AUROC from 0.665 prompted to 0.866, above 0.753 with STRIDE and Gloom training alone. In a separate four-round collection study, feedback-guided generation reaches 0.881 AUROC versus 0.770 for random collection at the same training-set size. Our trained monitors achieve Pareto-optimal cost–detection trade-offs among the evaluated configurations, with Qwen3.5-397B-A17B reaching detection performance comparable to Gemini 3.1 Pro at about 3.65-fold lower modeled inference cost. This framework connects failure diagnosis, data construction, and monitor training through controlled counterfactual pairs, providing a practical way to improve detection with agent monitors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.