Back to Basics: Formal Analysis Hints Improve Weak-to-Strong Monitoring
Abstract
Frontier agentic systems are capable of taking misaligned actions that may cause significant harm. Monitoring techniques are developed to oversee these systems in order to detect and prevent dangerous actions. However, current state-of-the-art monitors suffer from four persistent failure modes: complete detection failures, partial detections, accepting false benign framings, and suspicion score calibration errors. These failure cases can be efficiently combated by augmenting transcripts with supplementary hints, spanning command provenance, security-sensitive action flags, and task-disallowed effect markers. Provenance links allow for the recognition of diffuse attacks, while security and authorization flags simultaneously enable subtle attack recognition and provide an explicit reason for monitors to doubt the false benign framings they encounter, reducing rationalization and improving calibration. We find that augmenting weak extract-and-evaluate monitors with these hints significantly boosts performance with minimal computational cost. Furthermore, a failure mode analysis confirms that hints can remediate all four failure modes, notably almost entirely eliminating detection failures. A series of ablation studies emphasizes the efficiency and potential of the augmented extract-and-evaluate approach. Synthesizing our findings, we propose a tractable two-pronged approach toward enabling weak-to-strong monitoring.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.