Focused but Fragile: Diagnosing Tail-Rule Loss under Recursive Agent Skill Compression
Abstract
Compression can preserve an agent skill's routine while omitting the rule needed for an exception. We study this failure mode in an artifact-only maintenance regime, where each rewrite sees only its predecessor and cannot revisit the source evidence used to create the skill. We introduce TailSkills, a controlled benchmark of 208 task variants derived from 54 base tasks and 14 exception types. Each variant pairs an ordinary environment with an exception-triggering environment and is admitted only after reference-oracle execution and deterministic verification. The benchmark reveals a selective durability gap. Among the 56 variants with detectable registered markers, unit-weighted S4 retention is 48.7% for regular workflow units but only 15.7% for tail recovery units. A separate high-precision analysis over all 208 variants reproduces the artifact-level separation across four pipelines: at S4, regular retention ranges from 31.2% to 43.9%, whereas tail retention ranges from 2.1% to 15.2%. On a fixed set of 52 tail variants, S4 success is lower than S1 in three of four tested pipelines, with endpoint drops of 3.8–7.7 percentage points; the remaining pipeline follows a non-monotone trajectory and ends at its S1 level. Paired execution audits further show that artifact loss can become task failure in some executor–task settings, but does not produce a uniform downstream drop. These findings indicate that ordinary-task success is insufficient to assess the durability of compressed procedural skills: recovery rules require direct evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.