SwarmBlame: Content-Only Attribution for Adversarial Agent Swarms
Abstract
Coding agents increasingly work as swarms that edit one shared repository, so an operator who finds a planted bug or a covert message must determine which agent wrote it. Provenance records such as logs and version history can be rewritten by the suspects and detach from an artifact once it is copied, whereas a watermark puts the evidence into the text each agent generates, under a secret key. A core requirement of attribution is a framing bound, which states that an audit wrongly names an innocent agent only with a small, fixed probability. However, re-testing hundreds of watermark keys while an auditor reads file by file breaks this bound, a rewrite by another agent erases the author's mark, and one short artifact carries too little evidence. We propose SwarmBlame, an attribution protocol in which every agent samples under its own key with the standard KGW watermark and an anytime-valid test over all keys adds up evidence across files, naming an agent as soon as the evidence suffices while bounding framing for any reading order and stopping time. Composed keys record the chain of agents that rewrote a file, and seen-pair exclusion blocks framing by marks stolen from a victim's text. Across twenty models and six prior schemes, sixteen synthetic pull requests name their author in 99.9% of audits with 7 false accusations in 21,000, whereas re-testing detectors frame innocents in up to 8.93%, and optional learned weights nearly double code attribution without the prompt at audit time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.