FORTIS: Benchmarking Over-Privilege in Agent Skills
Abstract
Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This layer is widely treated as an organizational abstraction, but we argue it is also a safety boundary that current models routinely violate. We present FORTIS, a benchmark that evaluates agent skill safety across two stages: whether a model selects the minimally sufficient skill from a large, overlapping library, and whether it executes that skill without expanding into broader tools or actions than the skill permits. Across ten frontier models and three domains, we find that over-privileged behavior is the norm rather than the exception. Models consistently reach for higher-privilege skills and tools than the task requires, failing at both stages at rates that remain high even for the strongest available models. Failure is especially severe under the ordinary conditions of real user interaction: incomplete specification, convenience framing, and proximity to skill boundaries. The results indicate that the skill layer, far from containing agent behavior, is itself a primary source of unsafe escalation in current systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.