acceptodds
Under review as a conference paper at ICLR 2027

Did the Model Read the Paper? Attributing Paper-to-Code Scores to Their Sources

Abstract

Paper-to-code benchmarks are meant to measure whether models can turn research papers into working code, but their scores may also reflect information already available in the provided repository. We study this source-attribution problem in ResearchCodeBench, where only the target code span is masked while the other annotated code in the same file remains visible. We introduce paired interventions that separately remove the paper and this neighboring code while keeping the task and tests fixed. Across three models, removing the paper reduces pass rate by 1.6–8.1 percentage points (nominally significant for one model), whereas masking the neighboring code reduces it by 16.8–25.4 points and eliminates roughly two thirds of successful cases; most of this effect remains when the masked neighbors are replaced by a list of the names they bind. The asymmetry is concentrated in short targets nested inside another annotated snippet. A pre-specified split of the mask shows that most of it comes from the rest of that enclosing snippet, which the benchmark leaves visible, while the method's other components matter no more than the paper. The asymmetry also shrinks under the benchmark's line-weighted metric. We further find that naively masking multiple implementations makes the original prompt ambiguous, inflating the measured effect of code context. We fix this evaluation artifact by explicitly naming the target to restore, which recovers 6.5–8.6 points, while the larger effect of neighboring code remains. Our results show that repository-assisted paper-to-code benchmarks should control alternative implementation sources, starting with the code that encloses the target, before interpreting test-passing code as evidence that a model implemented the paper.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.