DR4Code-Bench: Benchmarking Research Experience Reuse Across Software Evolution
Abstract
Related changes in evolving software repositories often revisit the same concepts, creating opportunities to reuse earlier repository investigation. Whether such reuse improves coding success, reduces repeated investigation, or both requires direct evaluation. We introduce , an executable benchmark for evaluating pre-solution Investigation Report reuse across concept-linked software evolution tasks. It contains 88 executable tasks from 10 Python repositories organized into 22 concept-propagation chains. Each task is independently reconstructed and validated. Agents investigate the current repository before coding; only their Investigation Reports are carried across tasks, while prior agent patches, research traces, and oracle feedback are withheld. We evaluate six models on the 66 non-anchor tasks under No-History, Full-History, Recent-1, and Top-1 conditions, using a shared research-to-coding protocol and scoped outcome evaluator. Single-run correctness differences vary across models and policies. Three additional runs for DeepSeek-V4-Flash and Qwen3.8-Flash show that the direction of observed correctness gains and the strongest fixed policy can change across executions. In the six-model single-run evaluation, the reuse policies require 2.1–2.8 fewer research calls per episode on average across models and tasks, with small mean changes in coding calls. The reductions persist when comparisons are restricted to paired episodes in which research completes under both conditions. Repeated-run analyses show more consistent research-call reductions for DeepSeek than for Qwen. These results distinguish savings in repository investigation from improvements in final coding correctness and establish DR4Code-Bench as a setting for measuring both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.