Inspect or Roll Back? Asymmetric Diagnostics for Checkpointed Recovery
Abstract
When a stateful workflow fails, recent checkpoints may already contain the mistake. A controller can inspect an earlier prefix or discard it and repeat the remaining work. We study how the two kinds of inspection error should affect this choice: falsely trusting a bad prefix risks an incorrect commit, whereas falsely rejecting a good prefix wastes computation. An ordered recovery model separates these consequences and supports inexpensive asymmetric diagnostics. A three-cost representation accounts exactly for mistaken execution paths within a restricted policy class; we extend it to failure penalties and persistent hidden error modes. Experiments compare calibrated unit-test services, shallow lookahead, an edge-cutting baseline, and broader-history policies. In an executed multi-file code-edit workflow, asymmetric tests reduce the one-step controller's checking and replay work by 18%; two-step lookahead matches full compilation, and repairing the test oracle saves substantially more. Risk-matched comparisons and measured planning costs identify full compilation as a conditional refinement for costly, reused workflows. The main practical lesson is to improve and calibrate the diagnostic before using deeper recovery planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.