Patch-Bench: Evaluating Data and Privilege Minimization in Skill-Using Agents
Abstract
Agent skills encode vital domain knowledge and procedures to help agents generalize to unfamiliar tasks. As agents execute skills in real environments, these agents may access sensitive data or invoke system privileges that are unnecessary for the given task. Existing coding-agent benchmarks primarily evaluate whether an agent completes a task successfully, but do not assess whether it completes the task using only the minimal data and privileges required. To assess whether the minimization principle holds, we introduce Workflow Analyzer to identify potential privacy and security vulnerabilities in skill-conditioned coding agents and patch existing tasks to capture these overlooked risks through iterative task evolution. Following Workflow Analyzer, we present Patch-Bench, which contains 100 synthetic coding tasks constructed from existing skill-use tasks with functionally benign but security-relevant minimization requirements. Patch-Bench turns identified potential risks into reproducible and verifiable tests. For each task, we let agents run on Docker environments with independent utility and safety verifiers. This design distinguishes task competence from least-privilege execution. Patch-Bench enables systematic measurement of whether skill-using agents can preserve task utility while minimizing data exposure and operational authority. We conduct comprehensive evaluations on Patch-Bench to validate that existing benign skill-use agents may unintentionally request excessive data access and privileges out of the task scope. The results further suggest that functional success does not guarantee adherence to data-access and privilege boundaries during agent execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.