Delete or Preserve? An Empirical Study of the Effects of Deletion and Preservation Instructions on AI-Generated Code
Abstract
The two leading AI coding tools embed contradictory one-line instructions in their system prompts: Anthropic's Claude Code tells the model to "delete unused code completely," while OpenAI's Codex tells it to "never revert existing changes ... unless explicitly requested." Both ship to millions of developers; neither provider has published measurements of the effect. We test these instructions as a controlled manipulation across five LLMs (Claude Sonnet 4, Gemini 2.5 Pro, Qwen3-32B, Gemma 4-31B, and GPT-5-mini; n=10 per task per condition, 3,600 runs total) on 24 code repair tasks that each require a subtractive fix. The preservation instruction degrades code correctness on every model and by 13-15 percentage points on three of five (Claude p_Holm=0.006, Cliff's d=0.14; Gemini p_Holm=0.0003, d=0.14; Qwen3 p_Holm=0.004, d=0.13). The largest effect is on GPT-5-mini (the only OpenAI model in our panel, running its vendor's own prompt): a 39 pp correctness drop (p_Holm<10^-22, d=0.39). The effect survives Holm correction across the full 30-pairwise-test grid. The deletion instruction is neutral-to-helpful and never harms correctness. We complement this with an observational analysis of 57,371 commits from the AIDev dataset. OpenAI Codex, whose vendor ships a preservation instruction, has the lowest deletion ratio in deployed use, as the prompt-level mechanism predicts; strikingly, Claude Code deletes below the human rate as well, despite a vendor prompt that explicitly encourages deletion – evidence that base-model shyness can persist even against a pro-deletion instruction. GitHub Copilot and the autonomous agent Devin delete more than humans, while Cursor sits at the human deletion rate. We further analyze which additive failure modes the preservation prompt triggers – wrapping and downstream patching dominate – using a taxonomy of eight such modes as an analytical instrument. A follow-up experiment (Phase 2; three local models, 36 tasks per cell at n=5) tests six scope-limited variants against a three-part pre-registered gate (mean correctness recovery, over-deletion safety on benefit-measuring (BM) tasks – whose gold fixes preserve code – within 1 pp of blanket, and same-direction robustness on every panel model). All six candidates recover correctness on average (mean recovery 80-92%) and improve correctness over blanket on every model, but every one regresses BM safety on Llama 4 (+7 to +17 pp); exploratory per-model targeted prompts reproduce the same recovery-safety trade-off on Llama 4 and Qwen 3, so no instruction we tested is cross-model Pareto-improving over blanket preservation. The headline implication is that a single line in a system prompt imposes a measurable, replicable correctness penalty on the code produced by leading AI coding tools, and that the prompt's authors have a clean engineering choice with quantifiable consequences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.