When Small Evaluator Biases Become Permanent: Phase Transitions in Self-Evolving Agent Skill Libraries
Abstract
Persistent skill libraries allow large language model agents to reuse and refine executable artifacts. Recursive selection and retrieval, however, can amplify small evaluator preferences into lasting shifts in library composition. Identifying this amplification requires separating it from utility differences, repair asymmetries, retrieval preferences, and finite-population drift while modeling artifact evolution and endogenous exposure. We introduce the Causal Replicator–Retrieval model (CRR), a stochastic framework that tracks the quality-weighted frequency of a controlled, task-irrelevant attribute and the resulting cross-group exposure imbalance. CRR integrates quality-conditioned evaluator selection, directional transformation through mutation or repair, and retrieval reinforcement driven by accumulated exposure. Its mean-field dynamics characterize low-bias and biased equilibria, and its local Jacobian determines whether perturbations contract or amplify. We identify the underlying mechanisms using counterfactually matched skill variants, complete lineage and retrieval logs, and randomized channel-level interventions on evaluation, promotion, and repair. A hierarchical Bayesian state-space formulation estimates mechanism parameters and predicts held-out long-run regimes. The posterior informs a minimum stability-margin controller that selects evaluator calibration and exposure de-reinforcement while preserving task utility. We evaluate CRR in synthetic populations and executable-skill sandboxes through parameter recovery, unseen-phase prediction, channel-specific counterfactuals, and bias–utility–cost trade-offs. These evaluations provide a falsifiable framework for diagnosing and controlling evaluator bias in recursively updated skill libraries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.