Claimed, Shared, Verified: Auditing Data Availability in Quantum and Classical Computing Research
Abstract
Data availability is usually summarized as a single headline figure: the share of papers that release their data. Such a figure says how much data is shared, not what holds it there, nor whether stated availability translates into actual access. We bring two complementary approaches to this problem across quantum computing and classical computing research. The first is a large-scale automated classification of arXiv papers published since 2021, scoring each paper's use of external and self-generated data independently across six cases, ranging from directly accessible to claimed-but-unreachable. This lets us compare how the two fields actually share data, rather than how often they say they do. The second is a fine-grained verified-accessibility framework that distinguishes declared availability from independently confirmed access, recording not just whether data exist but whether they are locatable, complete, documented, and persistent, dimensions a binary availability measure cannot capture. Linking each paper to its publication venue, funding sources, and citation record, we test whether availability tracks venue policy, whether funder open-science mandates reach the work they cover, and whether shared data is repaid in citations. Together, the two approaches show that data availability is an outcome of the incentives around a paper rather than a property of the discipline it belongs to, and that declared availability systematically overstates verified accessibility. We release the paper-level classifications and linked metadata so these relationships can be tested further.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.