acceptodds
Under review as a conference paper at ICLR 2027

Do You Copy? When Language Models Overwrite Observed Values with Remembered Defaults

Abstract

Coding agents copy values from tool output into the files they write. We show that when the output names a well-known service on a non-default port, such as redis on 6392, language models overwrite the observed port with the default they remember. We call this failure default capture. Both values are true, so counterfactual knowledge-conflict benchmarks cannot measure it. We isolate it by renaming the service: the name effect is the real name's rate of writing the remembered default minus the rate under a made-up name of the same length. Under made-up names six instruction-tuned models copy the port from command output almost without exception, so a default written under the real name is attributable to the name. Default capture concentrates where agents meet it: returning command output as a tool result instead of a chat message triples the average effect, while authored configuration files stay near zero. In the preregistered chat test a docker ps row exceeds a .env line, and three of six models pass the per-model criterion. A header naming the source of the text, as in retrieval-augmented generation, lowers the effect below chat delivery, and under tool delivery dropping the service name from the line being written removes most of it. In a sandboxed agent loop, Qwen3.5-9B with thinking on sets the documented default in 24 percent of real-name episodes and 1 percent under made-up names, DeepSeek V4.1-Flash in under 1 percent of either. Yang (2026) found the same pull in Python import aliases; the six models repeat it, as a preregistered test predicted. Training toward the model's own behaviour under made-up names cuts the effect on a 7B model from 28.8 to 1.0 points under tool delivery on held-out services and layouts, with recall of remembered defaults within 3 points. Every repair trained only toward the page, and every inference-time method we compared that lowers the effect, trades copying the observed value against a second request: what port the service normally uses. Supervising that request too, on prompts rotated through raw, chat and tool form, lets a 1.5B model answer it at 76 to 83 percent in every delivery, against 31 to 47 for the frozen model, with the name effect near zero (exploratory).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.