When Failure Banks Kill Good Ideas: Calibrated Negative Memory for Autonomous Research
Abstract
Shared failure banks turn noisy experimental outcomes into reusable negative knowledge. We study when an irreversible binary failure record becomes a wrongful conviction: a genuinely improving idea is rejected and never retried. Under success-gated proposal supply we show theoretically that this local Type-II error can extinguish an entire research lineage—a deference phase transition and, at finite budget, an inverted-U relationship between deference and discovery; under exchangeable supply, sharing remains beneficial. Across two live ML loops and 122,817 benchmark convictions, wrongful convictions concentrate on marginal improvements, at a rate set by boundary density rather than effect variance; verification-stage near misses are 35–50% wrong and do not vanish as measurement noise shrinks. In controlled choice experiments, twelve deployed LLMs from nine families treat DEAD as a directive: they revisit 0/94 marked ideas and suppress an a-priori-best option despite identical evidence. Replacing binary records with evidence-graded, retryable statuses yields 30–32% more truth-validated discoveries in a controlled success-gated loop (8 replicates), while the advantage reverses in a low-noise, short-horizon cell. These results identify a conditional systems risk and a practical fix: make irreversible locking depend on recorded evidence and continuation value, with an explicit reopening path. The intervention is a schema change, not a model change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.