What Net Accuracy Hides: Auditing Repair and Breakage in Verify–Refine Workflows
Abstract
A small accuracy change can hide substantial changes in which problems an LLM workflow solves—and regressions may be attributed to refinement that never executed. We present an execution-path audit for verify–refine workflows that pairs baseline and workflow outcomes with the initial answer's source, verifier decisions, and calls admitted or refused by the budget. Across program synthesis and mathematics, these records distinguish answer regressions from the execution paths on which they occur. In a GPT-4o study, a code workflow changes accuracy by −3.3 percentage points (95% CI: [−8.7,+2.0]), while repairing 24.4% of baseline failures and breaking 13.8% of baseline successes. Under tighter budgets, 96 of 450 executions stop before any refinement runs: a wrong final answer in these runs cannot be attributed to an executed refinement. With equal refinement allowances and no blocked suffix calls, reusing a stored baseline output yields 5/347 breakages versus 27/347 for same-policy regeneration on one mathematics reference set. A second frozen reference set on the same tasks repeats the breakage ordering, but not the net-accuracy advantage. A serving control changes 8.2% of paired workflow outcomes under batching while net accuracy moves by only −0.22 percentage points. These findings support reporting repairs, breakages, and actual execution paths together, and evaluating preservation against the repairs it forgoes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.