1D-Wiki-Bench: Evaluating Knowledge Production and Consumption from Requirements to Code
Abstract
The value of reusable project knowledge depends on artifact quality and its utility across requirement clarification and code implementation. We introduce 1D-WIKI-BENCH, a dataset and evaluation framework that connects knowledge production, evidence use, and downstream development. It contains 133 tasks across 10 source groups spanning 12 repositories, with fixed sources, vague requests, reference specifications, and implementation criteria. Knowledge systems build artifacts before target requests are revealed. A clarification agent produces a specification for a separate implementation session; both stages reuse the artifacts while retaining raw-source access. We evaluate artifact quality, evidence acquisition and application, specification and patch rubrics, and stage-wise resource use. Raw-only achieves the highest specification and patch rubrics. On matched tasks, all six Wiki-producing methods improve implementation evidence acquisition by 2.25–10.95 percentage points over Raw-only, yet score 1.91–4.80 points lower on the patch rubric. In 71 OpenWiki task pairs, criteria linked to newly acquired implementation evidence gain credit overall, while other criteria lose more. Traces show direct evidence provision and navigation to original sources. OpenWiki and Graphify save 7.77M and 21.45M consumption tokens, respectively, but neither offsets its construction cost over the evaluated tasks. These results motivate evaluating knowledge production and consumption jointly through artifact quality, task fulfillment, and total reuse cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.