Closing the Text Channel: Typed Adaptive Evidence Acquisition for AI Control
Abstract
AI control uses trusted monitors to detect unsafe behavior from powerful, potentially adversarial AI systems. In code-generation settings, however, the material being inspected, such as source code, comments, and execution output, may itself contain attacker-written text that misleads the monitor. For adaptive monitors, this creates an additional vulnerability: such text can influence not only the final judgment, but also what the monitor chooses to inspect next, thereby shaping the evidence used for subsequent decisions. We introduce *Typed Adaptive Evidence Acquisition* (TAEA), which addresses this problem by placing both evidence acquisition and final scoring behind the same trusted, typed evidence boundary. Trusted probes convert program structure and behavior into predefined structured observations, preventing unrestricted attacker-controlled text from directly influencing either process. We prove that, for a frozen monitor under shared trusted randomness, edits that preserve all registered measurements leave the entire inspection path, score, and verdict unchanged. Within this protected interface, TAEA interprets evidence in context and selectively acquires additional measurements when they are expected to improve the decision. Experiments on two code-sabotage benchmarks show that TAEA achieves the highest among evaluated monitors and is substantially more robust under the evaluated adaptive attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.