acceptodds
Under review as a conference paper at ICLR 2027

V-TAP: Execution-Grounded Vector Guidance for Tool-Using Agent Red-Teaming

Abstract

Red-teaming tool-using language-model agents is expensive: testing a single candidate prompt can require multiple model responses and tool interactions. An important practical question is therefore how to uncover more safety failures within a limited evaluation budget. We introduce V-TAP, which uses the target model's internal response to a candidate prompt to decide whether that prompt merits a full agent execution. A scoring model learned from previous execution outcomes prioritizes candidates within Tree of Attacks with Pruning, while actual executions determine success. On 116 AgentHarm tasks with Qwen3.5-9B, V-TAP increases observed attack success from 27.59% to 29.31% under the same maximum rollout budget, a 6.25% relative increase in successfully attacked tasks. It succeeds on four tasks missed by random selection, while missing two tasks that random selection solves. Separately, a reproducibly fitted selector increases three-candidate coverage on 48 held-out pools from 16.67% to 18.75%, a 12.5% relative increase, reaching the successful-candidate coverage available in those pools. The online difference remains statistically inconclusive, and a 20-task Gemma-3-27B-IT evaluation does not reproduce the advantage. These findings motivate target-informed candidate selection as a way to improve the yield of costly agent evaluations, while identifying the evidence still needed to establish reliable gains across models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.