WITNESS: Benchmarking Dependency Analysis in LLM Agents Where Taint Tracking Ends
Abstract
The safe deployment of large language model (LLM) agents faces the threat of untrusted content delivered through prompts. Information flow tracking (IFT) is proposed for defense, but it requires estimating which inputs each output of the agent-driving LLM depends on, and no benchmark evaluates this estimation against human ground truth. We introduce Witness, a benchmark that treats an agent-driving LLM's output–input dependences as a standalone, testable oracle. Spanning eleven agents in diverse scenarios, Witness provides fine-grained human-annotated dependences, each typed as data or control, over both benign and prompt-injected trajectories. Evaluating twelve frontier and open-weight models, we find that they systematically miss true dependences: even the best model recovers an output's exact dependence set only half the time. This under-identification is the failure mode IFT cannot tolerate, since a missed dependence lets untrusted influence reach a sensitive action untracked.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.