acceptodds
Under review as a conference paper at ICLR 2027

Negative-Space Security: Omission Attacks and the Observability Bound on AI Monitoring

Abstract

Monitoring changes to a codebase or system typically means inspecting what was done: the diff that was produced or the commands that were run. We study omission attacks, where harm comes from a required action that was silently not performed. Because the omitted action leaves no trace in the diff or transcript, omissions expose an observability limit: if honest and attack trajectories induce the same visible-input distribution, any monitor of that input is at chance. We introduce OmissionBench, which combines 94 incident-anchored pull-request triples (with honest, commission, and omission arms) and 285 live-agent LinuxArena sysadmin trajectories. Diff-only LLM monitors score omissions 0.05–0.13 AUROC below matched commission attacks; a Sourcegraph Cody-style rubric proxy evaluated through Claude Sonnet reaches 0.65 AUROC on omissions. In live trajectories, the strongest transcript monitor falls from 0.96 to 0.76 AUROC as attacker capability increases, indicating that it detects execution residue rather than omitted obligations. Detection recovers when monitors read an independent obligation list or post-execution state: obligation-aware prompting improves omission AUROC by up to +0.16, and a simulated-output state-audit surrogate reaches 0.952 AUROC at 100% recall and 16% FPR. Robust AI-control monitoring should audit what agents leave behind, not only what they do.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.