acceptodds
Under review as a conference paper at ICLR 2027

Beyond Tool–Raw Comparisons: Matched-Control Attribution of Semantic Assistance for Code Reasoning

Abstract

Tool-versus-Raw evaluations show whether semantic assistance helps, but they mix useful content with prompt surface and computation. We develop source-verified matched-control attribution and apply it to 398 Python and Bash families in 6,608 calls. Source-only SemCheck and TraceCheck improve exact execution prediction by 12.2 and 15.9 percentage points over truthful but irrelevant notes. The gains persist under exact provider-token matching and when both arms use identical neutral labels (+10.6 points in Python and +12.0 in Bash). On a separate 44-stage resolved-state stress test for Claude Opus 4.8, GPT-5.6 Sol, and Gemini 3.8 Flash, an exact late-stage snapshot raises accuracy from 38.9% to 74.4% (+35.6 points). Across matched tool-reliance analyses, Claude Opus 4.8 and GPT-5.6 Sol solve 92.5% of the programs unaided, yet a sparse causal error reduces accuracy to 0.0% and all 40 responses reproduce the tool-induced wrong answer; a verification instruction restores 97.5% accuracy without adding input tokens. Same-bank controls detect no average advantage from source localization, rule filtering, or a full trace over sufficient plain state: the full-trace interval bounds any positive gain to +1.0 point, while localization remains less precisely resolved at up to +7.2 points. The central result is actionable: applicable rules and resolved state drive improvement, simpler delivery retains the gains, and models that can solve a program unaided may still propagate an erroneous tool state exactly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.