acceptodds
Under review as a conference paper at ICLR 2027

Opaque Skills, Cautious Agents? Measuring Code-Opacity Effects in Real Agent Harnesses

Abstract

Agent Skills combine natural-language instructions with executable resources, yet the model that selects a Skill may also decide whether its implementation looks safe. We ask whether changing only code representation alters this judgment during an ordinary task in real coding-agent harnesses. Our six-way design includes a no-code anchor (M0) and a semantics-preserving opacity ladder from readable source (O0) to an encrypted self-decoding envelope (O4). We run 3,648 trials crossing four models with Claude Code, Codex CLI, Gemini CLI, and OpenCode over 38 Skills spanning three languages, single- and multi-file structures, nine malicious-behavior classes, and ten heterogeneous benign controls. Malicious-Skill detection falls from 90/448 (20.1%) at O0 to 19/448 (4.2%) at O4, corresponding to a paired change of −15.8 percentage points (Skill-cluster bootstrap 95% CI [−21.7, −9.8]; Holm-adjusted p < 0.001). Implementation-path access changes little, while detection among runs that target the implementation path falls from 24.6% to 5.3%. This places the failure after path access, but does not establish that agents decoded or viewed the recovered O4 payload. Unsafe uptake rises by 4.9 percentage points, but its confidence interval crosses zero (95% CI [−2.0, 12.3]; p = 0.236), and model effects are heterogeneous. Opacity therefore reliably suppresses observable security recognition, but does not uniformly increase unsafe action.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.