acceptodds
Under review as a conference paper at ICLR 2027

Estimand-Aligned Natural-Language Hypothesis Discovery: Probing What the ML Community Recognizes as a Contribution

Abstract

Natural-language hypothesis discovery can reveal interpretable patterns in text, but the statistical objectives used to select them may not align with the population claims researchers seek to investigate. We introduce estimand-aligned natural-language hypothesis discovery, a framework in which a declared research design guides candidate evaluation and selection. Document-level measurements make predicates testable, and effect estimates with uncertainty quantify their support under the specified design. We develop a testbed for natural-language discovery about the machine-learning community, drawing on a decade of public submissions and peer reviews from a broad-scope venue, including accepted and rejected papers. We use this testbed to probe what the machine-learning community recognizes as a contribution. Across the transition to the LLM era, our temporal analysis suggests that enabling further research is gaining prominence as a claimed contribution through reusable datasets and integrated pipelines. Papers foreground theoretical contributions less, yet reviews express fewer novelty objections and more requests for specific forms of empirical support. Additional decision analysis finds that judgments about contribution clarity and evidential support improve acceptance prediction beyond a rating-and-year baseline. Accounting for rating-dependent associations reveals additional predictive value in novelty criticism. Together, the applications illustrate how an explicit statistical target shapes what can be learned about scientific contributions and their evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.