Beyond Robust Scoring: A Four-Stage Audit of Persistent Skill Writes
Abstract
Persistent skill writing converts textual advice into parameter updates that remain active after the text is removed. We show why robust scoring alone is insufficient and turn the resulting audit into a write-control method. The audit separates score comparability, candidate support, causal realization, and behavioral-label validity. Its findings motivate Audit-Guided Write Control (AGWC), which combines support-aware stopping, paired shadow updates, context-dependent re-evaluation, and set-level commit/rollback. On ALFWorld, the reported 288-candidate, five-pool evaluation gives AGWC 44.48% scenario-OOD success versus 36.52% for Direct-Validation, while behavioral regression decreases from 5.10% to 3.60%. With 18 selected skills per method, the advantage remains 6.60 percentage points, despite identical observed proxy harm. Shadow probes predict independent full-write effects with AUROC 0.74 and Spearman correlation 0.48. After fresh retraining on disjoint evaluation games, held-out success improves from 22.22% to 28.02%. These aggregate results support persistent writing as a sequential verification-and-commit problem: comparable scoring proposes interventions, while explicit behavioral checks govern deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.