Not Every Failure Needs a Patch: PAUSE for Selective Program Repair under Uncertain Execution Feedback
Abstract
LLM-based automated program repair increasingly relies on iterative execution feedback to guide patch generation. However, systems often treat each failure as evidence that the program should be modified. Because test and runtime signals may be incomplete, noisy, or environmentally induced, this assumption can cause unnecessary patches, ineffective retries, and unreliable termination. We present PAUSE, a selective program repair framework that determines whether modification is warranted, whether the current candidate is acceptable, and whether continued repair remains worthwhile. PAUSE coordinates specialized agents to analyze repair context, reason over execution evidence, decide between Patch and NoOp, and learn from repair trajectories. Its dual-condition governor separately evaluates candidate acceptability and continuation value. Informative trajectories are distilled into reusable decision constraints rather than patch templates. Experiments on three benchmarks spanning Python and Java and dedicated control cases show that PAUSE improves confirmed repair rate by 2.5–3.0 percentage points over the strongest controlled baseline, limits over-repair and unsupported acceptance to 1.7% and 2.0%, respectively, and reduces model calls by 44.6% relative to fixed-budget execution. These results demonstrate the value of selective action and trajectory-aware governance for LLM-based program repair.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.