Ephemeral Models, Durable Claims: Closed-Model Dependence and the Preservation of AI Research
Abstract
Large language models (LLMs) are routinely used inside modern AI experiments: to generate or annotate data, to provide comparison points, and increasingly to evaluate other systems. When the exact model artifact can be retained locally, later researchers can often preserve that dependency. When access is mediated only through a provider-controlled service, they usually cannot. Using an end-to-end measurement pipeline validated on a blinded random sample (F1=0.967), we measure this dependence across 53,538 final archival papers from ACL, EMNLP, ICLR, ICML, and NeurIPS (2023–2026) and assess whether the referenced model artifacts persist. We find that 17,841 papers (33.32%) contain a covered closed model in at least one scientific role. Across the common five-venue window, prevalence rises from 14.96% in 2023 to 37.98% in 2025, while use of closed models as judges rises from 1.47% to 11.50%. At a September 24, 2026 lifecycle cutoff, seven of the 20 most frequently used normalized models are already retired; several others are deprecated, on explicit snapshot-shutdown paths, or named too coarsely to identify a unique served artifact. We characterize this exposure and derive practical preservation norms for durable scientific evidence, spanning model reporting, peer review, archival traces, and post-commercial research access.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.