Improving scalable oversight with co-trained monitors
Abstract
Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by _co-training the monitor alongside the worker_, and explore both supervised and self-supervised approaches. The supervised approach involves oracle access to gold-standard judgements from a principal (e.g. human). Monitoring – with vanishing error and query rates – turns out to be possible exactly when the class of worker strategies has finite Littlestone dimension, connecting our setting with an established literature on adversarial online learning. We also propose a self-supervised co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate higher-quality labels, then trains its standard-compute policy on those labels. We analyze conditions under which this procedure provably preserves the intended monitoring objective, keeping up with worker-induced distribution shift. We stress-test both approaches in code-security settings where the worker is trained adversarially to fool the monitor. Our results suggest that adaptive monitors are better at keeping pace with evolving worker strategies, while fixed monitors are more vulnerable to evasion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.