acceptodds
Under review as a conference paper at ICLR 2027

WITNESS: Benchmarking Dependency Analysis in LLM Agents Where Taint Tracking Ends

Abstract

The safe deployment of large language model (LLM) agents faces the threat of untrusted content delivered through prompts. Information flow tracking (IFT) is proposed for defense, but it requires estimating which inputs each output of the agent-driving LLM depends on, and no benchmark evaluates this estimation against human ground truth. We introduce Witness, a benchmark that treats an agent-driving LLM's output–input dependences as a standalone, testable oracle. Spanning eleven agents in diverse scenarios, Witness provides fine-grained human-annotated dependences, each typed as data or control, over both benign and prompt-injected trajectories. Evaluating twelve frontier and open-weight models, we find that they systematically miss true dependences: even the best model recovers an output's exact dependence set only half the time. This under-identification is the failure mode IFT cannot tolerate, since a missed dependence lets untrusted influence reach a sensitive action untracked.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.