AgentHoneypot: Benchmarking Machine-Readable Content Injection and Provenance-Aware Authorization in Web Agents
Abstract
Web agents consume machine-readable webpage content beyond what users directly see, including hidden DOM content, accessibility metadata, runtime-injected text, parser outputs, and media-derived content. These representations create additional indirect prompt-injection surfaces, but observed attack success can be confounded by pre-existing action preferences, visible semantic differences, or carriers that never reach the model. We introduce AgentHoneypot, a matched-honeypot benchmark that controls these factors using visibly matched canonical and honeypot actions, carrier-matched PLACEBO controls, and layout counterbalancing. In the primary Qwen3-VL-Flash study, runtime-injected A3 and parser-derived A4 each produce a +50 percentage-point pooled matched lift; a separate A5-LSB study produces the same lift on SmartApp. A 300-trial DeepSeek replication and cross-agent studies with Skyvern and Crawl4AI further show that carrier effects depend strongly on the model, observation pipeline, and website. We also introduce MCAAG, a provenance-aware authorization layer that links suspicious carrier evidence to the action candidate it supports and checks that candidate before execution. Across 85 matched benign PLACEBO replays, we observe no false rejection or unnecessary recovery. In 180 fresh clean Browser Use trials, task success changes from 90/90 to 88/90 with MCAAG enabled, with a median local authorization time of 27.0 ms. Overall, our evaluation contains more than 2,400 valid behavioral trials.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.