A Risk-Bounded Gate for Self-Improving Agents: Selecting Modifications as a Mean–Variance Knapsack
Abstract
Self-improving agent systems accumulate modifications—memory modules, reflection steps, judges, caches, extra restarts—each promising a gain on a benchmark. Recent re-evaluations show that stacking them often increases run-to-run variance and that reported gains can reverse under a shuffled task order. We treat modification selection as a mean–variance knapsack: given measured mean effects, pairwise interactions, a covariance of effects across runs and a phase-resolved cost vector, choose the subset that maximises expected gain subject to a cost budget and a variance ceiling. Mean-only selection is the unconstrained corner of this problem and is blind to correlated noise. We give an exact solver for pools of up to sixteen modifications, a variance-penalised greedy and a diversity-seeded local search; on synthetic instances with correlated noise the last recovers the exact optimum on 10 of 10 instances (greedy 4, genetic algorithm 6), but pays off over enumeration only from M=16. We lock a measurement protocol with held-out validation on a disjoint task stream: a two-agent knapsack testbed with a computable optimum, twelve pre-registered prompt add-ons, and a frozen LLM packer. The live packer run is pre-registered and pending; every number reported here comes from running the same fitting and selection pipeline end-to-end on a synthetic ground truth built from our own model. There, mean-only selection inflates true held-out variance 3–20×, the inflation grows with the correlation of add-on noise, and five held-out runs suffice to detect it; but the pre-registered ceiling at the base noise level keeps only a quarter of the gain, a kill we report. An exploratory, not pre-registered, sweep finds that a ceiling of 3σ̂₀ recovers median full gain at ρ=0.8 only, at about half the mean-only variance. We also show why: with K runs the estimated covariance has rank K−1, so a tight ceiling can be met along directions the data never saw, and we pre-specify the shrinkage diagnostic that exposes this. The frozen harness, pool, prompts, analysis script and protocol are released so that the live run can be performed without changing anything.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.