SkillVigil: Evidence-Supervised Guardrails Against Skill Injection in LLM Agents
Abstract
Coding agents now install third-party skills: bundles of instructions and executable scripts that enter the agent's context and run with its privileges. A poisoned bundle does not act on its own but recruits the agent that loads it, and it is reloaded every time the skill is used. The guardrails deployed for this surface check a bundle before it is loaded and stop there. Runtime monitoring is a separate line that watches the execution, rarely attributing it to the bundle that introduced it. Either way the operator learns at most that something is wrong, not which step is implicated or what text implicates it. Scored the same way on the same inputs, three of six general-purpose guardrails fall within of chance on skill bundles: a bundle reads as ordinary documentation until one line invokes a bundled script. We train a single stage-conditioned adapter that screens the bundle before execution and audits the resulting trace afterwards, supervised to return a citation and an explicit abstention alongside the verdict. It reaches AUC on MalSkillBench, where the best of those six reaches . It names the implicated call on of execution windows and quotes it on , whereas GPT-4o and Qwen3.8-B prompted zero-shot reach and on the locator, a gap of training exposure rather than capability. Adding the channels one at a time under a fixed budget moves verdict AUC by at most in paired estimates on the same executions and bundles, and makes abstention, otherwise ranging from to across backbones, uniform. Localization succeeds where the answer can be copied from the input and degrades where it must be counted to, a limit we trace to the output format: probes read from the same hidden states the position it still fails to report.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.