The Review Tax: From Automatic Review to Post-Edit Control in Coding Agents
Abstract
After a coding agent produces an edit that can be applied, should the harness automatically spend another model call reviewing it? We study this _**review tax**_: computation committed to review simply because a patch exists, before there is evidence that review is worthwhile. In a controlled three-call harness on 60 SWE-bench Verified tasks under the same 32k context limit, review consumes 20.8% and 23.7% of all model tokens for two Qwen coder configurations. More than 83% of those review tokens are spent reading the input. Measuring the return requires care, since separately rerun review and no-review arms can produce different patches before review begins. When we instead fork the same saved patch on 200 tasks at 64k, review gains 5.0 percentage points, with a 95% interval of , while a prompt variant gives a smaller estimate whose interval includes zero. These findings motivate treating the stage after an edit as a decision. In a separate sequential campaign with 308 task entries from 305 SWE-bench-Live issues, a learned policy resolves 6.5 points more than a state-blind schedule, at higher reported cost. A further repetition again favors learned control over the schedule but leaves its advantage over a correctness gate unresolved. When a public-test result has already been obtained and charged for, showing it to later decisions yields 15.5 points higher resolution than masking it. An applicable edit is therefore an opportunity to review. Whether to do so should depend on the candidate patch, the evidence already available and the cost of continuing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.