acceptodds
Under review as a conference paper at ICLR 2027

SkillFence: Execution-Grounded Trust Boundaries for Skill-Augmented LLM Agents

Abstract

Skill-augmented LLM agents increasingly act on externally supplied workflow artifacts, such as skill files and tool outputs, that are operationally relevant but may be attacker-controlled. We identify skill-context injection, in which such artifacts are elevated from contextual evidence to execution authority, so that forged approvals, routing notes, or policy claims become unsafe tool actions. To measure this threat, we build a black-box benchmark that rewrites 28 seed attacks into enterprise-style skill and tool artifacts using safety-ablated open-weight models. On agents built on GPT-5.4, Claude Sonnet 4.6, and Kimi-2.5, these artifacts achieve broad attack success rates (Broad ASR, which counts unsafe actions, authority confusion, and reasoning contamination) of 38.39%–76.07%. We propose execution authority separation: artifacts may inform a task, but only a trusted authority root, such as the user or system policy, may authorize a sensitive action. SkillFence enforces this principle by extracting authority claims from artifacts, estimating residual authority confusion, and gating each sensitive tool call against the trusted root. It reduces Broad ASR from 70.71% to 6.96%, from 55.00% to 5.18%, and from 41.16% to 2.86% on the three backbones while retaining 98.67%–100% benign utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.