Catching Schemers in the Act: Training Online Monitors to Detect Covert Misalignment
Abstract
As LLM agents work longer with less supervision, oversight increasingly depends on monitors to detect _scheming_. Prior work typically employs scheming monitors that run asynchronously after the trajectory has completed, at which point harm may be irreversible. We build _online_ black-box, action-only monitors that assess trajectories as they unfold. For this purpose, we construct a pipeline to augment datasets with per-step labels, and introduce _timely_ metrics that credit a flag only while the harm can still be averted. Additionally, an online monitor must be accurate, cheap, and fast: prompted frontier monitors detect well but are costly and slow, while smaller prompted models are cheap but miss critical information. We train open-weight monitors, several of which are Pareto-optimal on both cost-performance and latency-performance curves. Our trained Qwen3.8-27B outperforms GPT-5.6 Sol at lower cost and half the latency. It nearly matches Gemini 3.7 Flash and Claude Opus 4.8 in detection performance at and lower cost while being slightly faster than both. Only Gemini 3.1 Pro is clearly stronger, at the cost and the latency. We further analyze the relationship between the composition of the training mixture and OOD transfer, and find that many residual failures come from monitors misweighing evidence they have seen, rather than missing it. Together, our per-step labels, metrics, and trained monitors offer a practical recipe for low-cost, low-latency oversight of agents as they act.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.