When and Why Do Distributed Attacks Evade Code Monitors?
Abstract
When an untrusted coding agent works on a repository with a series of pull requests (PRs), work for a hidden side task can be distributed across PRs. Spreading the attack may or may not help against a monitor that reviews each PR's diff, depending on the side task. We define criteria for decomposable side tasks that in theory allow an attacker to do every component more innocently. We build a dataset of decomposable side tasks. Across more than 11,000 trajectories from four attacker models and eight monitor models, gradual attacks are generally more evasive than the best single-PR placement on the cache-poisoning side tasks (43 of 252 comparisons are significant after Holm correction, none in the other direction), while single-PR attacks tend to be more evasive on the credential-exfiltration side tasks. We study detectability of these attack types both through what the monitor sees and judges, and through activation probing. The most relevant factors are how well the component fits the scope of the current task, if the monitor checks diffs against the task's requirements, how much suspicion honest code raises, and connections of current changes to earlier changes. Linear probes on the attacker's residual stream predict its next unauthorized action about equally well under both strategies. We release the dataset and code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.