acceptodds
Under review as a conference paper at ICLR 2027

ReSpec: Learning Paper Reproduction via Synthetic Supervision and Process Rewards

Abstract

Faithful paper-to-code reproduction requires translating a paper’s technical specifications into an implementation. Learning this capability requires intermediate supervision and fidelity feedback, which paper-repository pairs alone do not provide. We present ReSpec, a framework that automatically constructs synthetic supervision and checklist-based process rewards from papers. Its data engine selects 6,172 paper-repository pairs through implementation-fidelity assessment and synthesizes 430,428 examples for six reproduction sub-tasks. It also constructs paper-grounded checklists to define fine-grained rewards for the plan pipeline and file-scoped code writing, yielding 76,448 samples for parallel training through task-decomposed reinforcement learning. Using these paper-derived signals, we train the ReSpec model through supervised fine-tuning followed by reinforcement learning. On PaperBench Code-Dev, the ReSpec model equipped with the DeepCode-Base scaffold and two-pass plan refinement outperforms Qwen3.5-35B-A3B in the DeepCode-Base scaffold by 5.2 points (70.6 vs. 65.4) with one-sixth as many active parameters (0.5B), while approaching Claude-Sonnet-4.5 (73.5) and GLM-4.7 (75.0). These results demonstrate that automatically constructed process supervision and rewards can turn paper-to-code reproduction into a learned capability in a compact specialist.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.