acceptodds
Under review as a conference paper at ICLR 2027

SkillGovernor: Localizing and Governing Risk across Agent Skill Graphs

Abstract

LLM agents accomplish complex real-world tasks by loading skills, which turns the skill package into an exploitable control surface. Attackers manipulate the content, structure, and interrelationships of skill files to inject payloads, so the skill still appears benign while the payloads couple with the procedures the agent depends on. Existing audit methods locate risk without modeling how it propagates across the skill's structure, and runtime guardrails block risky actions while leaving the risk-bearing content in place. We propose SkillGovernor, an end-to-end framework that governs a risk-bearing skill at its source. Its Risk Audit organizes the skill into a graph, reasons over it in cascade, and joins scattered risk into a risk-connected view. Its Dynamic Governor operates on that view and chooses the least intervention among graded operations, one that removes the risk and keeps the skill's legitimate function. We evaluate on 78 cases from two public skill-security benchmarks, across five configurations on three agent systems. SkillGovernor reduces the attack success rate (ASR) by 26 percentage points on average while maintaining task completion, reaching the lowest or tied-lowest ASR in every configuration. Its structure-based audit raises the injection point detection rate from 0.26 for a naive judge to 0.76. Together, these results show that a skill's risk can be governed at its source without compromising functionality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.