BrowseSafe++: Benchmarking Prompt Injection Detection Across Web Contexts, Languages, and Browser Views
Abstract
Browser agents need to make decisions and execute actions on behalf of users while consuming untrusted webpages. This creates the chance for malicious instructions embedded in webpages to redirect legitimate tasks. Existing benchmarks establish this threat, but broader evaluation requires consistent coverage of the web domains, languages, and page views encountered in deployment. We introduce BrowseSafe++, extending BrowseSafe from 14,719 to 38,675 entries. We start with prompt injections observed in the wild from production website captures, preserve the original injections, and rewrite sensitive portions to protect privacy. Each entry pairs an attacked page with a control page that differs only by the injection, provides HTML, Markdown, accessibility tree, and screenshot views, and is linked to a legitimate task and expected outcome. The benchmark contains twelve domains, twenty-four attack types, and eighteen placements. Using these pairs as supervision, we train BrowseSafe-Guard, compact 0.8B text and multi-modal detectors that outperform frontier models while raising far fewer false alarms. We also introduce BrowseSafe-Adapt, which refines benchmark injections from browser-agent feedback and shows that attacks succeeding on a few percent of tasks when replayed succeed on up to 95% after adaptation. The benchmark, model, and SDK for adaptive framework will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.