acceptodds
Under review as a conference paper at ICLR 2027

When Code Changes, Answers Don't: Breaking Semantic Inertia in Code Models

Abstract

Code models increasingly reason about programs that change, yet their predictions often remain anchored to what the program did before an edit. We call this failure semantic inertia. A small edit may change the output for one input while preserving it for another, so reliable reasoning requires a selective update rather than a blanket reaction to modified code. Across three execution-tuned Qwen3 models, 93% of edited-program errors following a correct original prediction repeat the original output. On a separate frozen panel of 15 models and checkpoints, semantic inertia spans model families and remains measurable even in the strongest hosted models. We introduce the Semantic Inertia Benchmark to measure this capability directly. Each group pairs an original and edited program with a Changed input that requires Revision and an Unchanged input that requires Preservation; Joint accuracy requires all four program–input predictions to be correct. We then propose Counterfactual Regrounding, which trains on these execution contrasts, emphasizes edited-program answers, and retains correct original-program judgments. Inference remains a single forward pass on one program and input. On 1,000 held-out APPS-derived groups, Counterfactual Regrounding raises Qwen3-4B's Joint accuracy from 9.5% to 16.0%, lowers inertia from 75.4% to 63.5%, and improves Preservation. Its gains grow with Qwen3 model size, transfer to OpenCoder-1.5B, extend to unseen edit families and open-ended generation, and persist on the separate 2,323-group panel. These results establish selective semantic updating as a distinct capability—and execution contrasts as an effective way to train it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.