acceptodds
Under review as a conference paper at ICLR 2027

ATTRACK: Attribution-based Taint Tracking with Provable Bounds for Agent Security

Abstract

Language-model agents act on web pages, emails, and files that may contain indirect prompt injections: instructions planted to redirect agents toward an attacker's goals. We propose ATTRACK, a defence that measures how strongly untrusted content influences each proposed tool call. Its influence test deletes untrusted content and scores the same call again, blocking calls whose log-probability drops beyond a threshold. Citation checks, which detect instructions attributed to untrusted sources, stop most attacks early. The influence test is a certified backstop even when the citation checks are deceived: it bounds the probability that the agent takes a malicious action and the defence misses it, relative to its probability under trusted context alone. Across four AgentDojo suites and two models, ATTRACK blocks every observed attack on 194 matched cases per model. Without attacks, it keeps 89% of the undefended agent's task utility on Qwen3.8-27B and 79% on Gemma-4-26B, compared with 51% and 54% for provenance-taint isolation. On Gemma, by contrast, three detectors (spotlighting, MELON, and a prompt-injection classifier) allow attacks in up to 12.9% of cases. An adaptive attacker finds no successful attack against ATTRACK in 832 payloads over 16 Gemma cases, but finds new ones against all three detectors. For half of AgentDojo's attack cases on Qwen and most on Gemma, the bound keeps the chance that the attacker's exact goal call gets past the defence below one in ten thousand.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.