Are Agents Leaving Your Code Stupid? Predictive Refactoring for Post-Repair Hardening
Abstract
Software engineering agents are evaluated mainly by issue-resolution accuracy, and their workflows are reactive: the agent fixes the reported bug and stops, even when repository history marks the surrounding code as defect-prone. In tangled, high-churn legacy code, a patch can then pass the benchmark while leaving brittle interfaces, recurring defect hotspots, and hidden temporal coupling in place. We propose Memory-Augmented Predictive Refactoring (MAPR), a two-phase workflow that treats a solved repair as the starting point for regression-aware hardening. MAPR builds repository memory from historical pull requests, summarizing bug density, author-linked defect patterns, temporal coupling, review coverage, and code churn. After Phase 1 resolves the issue, Phase 2 replays the Phase 1 patch in a fresh sandbox and hardens the repaired code under a test-driven RED–GREEN–VERIFY protocol that targets guard clauses, diagnostic error messages, and a controlled blast radius. We evaluate MAPR on Phase 1-solved SWE-bench Verified instances with four backbones, together with a 348-instance MiniMax-M2.5 run over the full benchmark. Full Phase 2 raises guard density by a factor of 4.5–7.0 on every backbone and keeps 96.4–100% of repairs on the three backbones with complete harness logs; on the 348-instance run, it raises guard clauses from 68 to 798 while keeping 96.5% of repairs. The RED–GREEN–VERIFY scaffold accounts for most of the gain in hardening quality: removing it lowers the share of error messages that report runtime state on every backbone, for example from 79.7% to 30.4% on MiniMax-M2.5. Separating the two stages also matters: a single-session agent with the same scaffold and step budget loses 6–14 repairs on these backbones, versus 0–2 for MAPR. Falling back to the Phase 1 patch whenever Phase 2 breaks a repair restores every repair and retains at least 97% of the added guards. The anonymized code, data, and analysis artifacts are available at https://figshare.com/s/b1f96ed457df4694a5a4.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.