acceptodds
Under review as a conference paper at ICLR 2027

What Net Accuracy Hides: Auditing Repair and Breakage in Verify–Refine Workflows

Abstract

A small accuracy change can hide substantial changes in which problems an LLM workflow solves—and regressions may be attributed to refinement that never executed. We present an execution-path audit for verify–refine workflows that pairs baseline and workflow outcomes with the initial answer's source, verifier decisions, and calls admitted or refused by the budget. Across program synthesis and mathematics, these records distinguish answer regressions from the execution paths on which they occur. In a GPT-4o study, a code workflow changes accuracy by −3.3 percentage points (95% CI: [−8.7,+2.0]), while repairing 24.4% of baseline failures and breaking 13.8% of baseline successes. Under tighter budgets, 96 of 450 executions stop before any refinement runs: a wrong final answer in these runs cannot be attributed to an executed refinement. With equal refinement allowances and no blocked suffix calls, reusing a stored baseline output yields 5/347 breakages versus 27/347 for same-policy regeneration on one mathematics reference set. A second frozen reference set on the same tasks repeats the breakage ordering, but not the net-accuracy advantage. A serving control changes 8.2% of paired workflow outcomes under batching while net accuracy moves by only −0.22 percentage points. These findings support reporting repairs, breakages, and actual execution paths together, and evaluating preservation against the repairs it forgoes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.