VeriLoop Coder-E1: Toward Verifiable Recursive Self-Improvement through the VeriLoop Spiral
Abstract
Self-modifying coding agents may retain locally successful changes without establishing whether those changes improve subsequent error correction. We present VeriLoop, an evidence-governed framework that separates within-task verification from cross-task method inheritance. VeriLoop organizes the improvement process into six operations: evidence, falsification, inquiry, revision, verification, and verified internalization. Its dual-gate protocol first validates a correction within its declared scope and then conditions inheritance on subsequent independent evidence, while preserving provenance and explicit revocation paths. We instantiate this framework as VeriLoop Coder-E1, built on a frozen Qwen3.6-27B backbone with four detachable Surface Host adapters and an external Self-Harness. VeriLoop achieves 85.20% on SWE-bench Verified, 62.38% on SWE-bench Pro, 76.40% on Terminal-Bench 2.0, and 33.63% on DeepSWE. Three classification adapters improve performance on held-out capability endpoints, while a fourth provides uncertainty estimates. Their LoRA matrices contain 3,481,640 parameters, corresponding to 0.012895% of the backbone. Mechanism audits across 115 episodes document selective admission and intact control boundaries. Together, these results support the feasibility of a practical and reversible method-inheritance architecture, while not establishing that inheritance causally improves downstream task success.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.