Learning to Repair the Learner: Evaluating Recursive Improvement in Neuro-Symbolic World Models
Abstract
Self-improving agents are usually credited through descendant-performance curves: later lineage members outscore their ancestors. Such curves misattribute improvement even in a fully inspectable system. We formalize a checkpoint separating world model, repair procedure, evidence and a fixed kernel, and derive a ladder of matched interventions, from crossed model/procedure states and owned versus replayed evidence to a code×history crossing, a generator lock and non-recursive baselines. We apply it to a neuro-symbolic lineage that edits its repair procedure and successor generator, on an original cohort and replications with 252 fresh families declared before execution. In replication, inherited procedure explains 45% of the naive descendant gain; model structure explains the rest. Under a shared source history, continued editing detectably improves on freezing, but through final code (history effect ≈ 0), and is practically equivalent to choosing one fixed behavior per family on the same continuation evidence with fewer fits; generator editing is practically equivalent to locking it. A positive control shows the ladder detects planted generator and history value. Because the base generator indexes proposals by history length, sharing a history between source lineages decides whether checkpoints lock into single-term repair (504/504 versus 75/504); with separate histories the gain over freezing is no longer detected. A replicated gain thus hinges on the checkpoint’s history contract, the failure our formalism exists to expose.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.