Causal Promotion Gates for Secure Self-Evolving Agent Skills
Abstract
Persistent skills enable large language model agents to reuse procedures distilled from prior executions. Promoting candidate skill updates to a shared repository requires determining whether persistent encoding improves performance on independent tasks while avoiding safety and subgroup regressions. Existing promotion mechanisms rely on recurrence, model judgments, source-trajectory attribution, or aggregate validation, but do not isolate the effect of persistently encoding a candidate rule. We introduce the Causal Promotion Gate (CPG), which formulates skill promotion as a distribution-level paired causal decision over persistent policy updates. Given a frozen agent policy, a current skill, and an atomic candidate update, CPG treats the skill version as the intervention and compares matched executions with identical task instances and initial states. The protocol excludes validation tasks that share provenance with the evidence used to generate the update, controls reusable randomness, and randomizes treatment order to reduce execution artifacts. The paired outcomes estimate utility, safety, and pre-registered group-specific effects under a fixed validation budget. One-sided confidence bounds, minimum-effect thresholds, and worst-group constraints yield an auditable three-way decision: promote updates with reproducible benefits and acceptable risk, reject clearly harmful updates, and defer updates when the evidence or protocol validity is insufficient. We evaluate promotion accuracy, calibration, deployment effects, heterogeneous harm, robustness, and validation cost on planted-causal benchmarks and through integrations with trajectory-to-skill pipelines. We compare CPG with common promotion and attribution strategies under matched validation budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.