When Do Independent Scientific Alternatives Emerge? Evaluating LLM Predictions of Waiting Times
Abstract
The wait from a scientific contribution to its first independent, functionally equivalent alternative describes when another route to the same capability becomes available. We examine whether evaluation scores reflect how well large language models rank these reinvention times. In a post-hoc audit of an author-built dataset and saved predictions from three models, we trace waiting-time labels to their target works and decompose the ranking score by comparison type. At least 8 of 34 eligible observed held-out examples have labels inconsistent with the target work's interval under the dataset's own rules. Adding curator-specified lower bounds (claims of no alternative before a cutoff) raises the C-index ranking score from 0.42–0.45 to 0.61–0.67 on the models' respective valid subsets, without changing any observed-interval prediction ordering. The increase comes from comparisons between observed intervals and lower bounds. A diagnostic using only record type scores higher still, without ranking waiting times within either type. The score increase persists after recomputing the known mismatches from retained records, including on samples valid for all three models. Neither excluding the known mismatches nor recomputing their labels establishes above-chance observed-interval ranking. Two models exceed the specified output-failure threshold, so their affected results remain exploratory. This case shows why assessing waiting-time ranking requires checking both what each label measures and which comparisons produce the score.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.