acceptodds
Under review as a conference paper at ICLR 2027

Beyond Model Leaderboards: Separating Dataset Design from Model Capability in Clinical Prediction

Abstract

Clinical prediction benchmarks typically fix a dataset and compare models, leaving the consequences of dataset construction underexamined. We introduce a self-contained dataset design system (SDDS) that separates design objects—source, population, features, missingness, and longitudinal structure—from confounding objects governing the model, task, preprocessing, and execution conditions. Individual-object studies and joint experiments over representative factors distinguish object impacts, within-object factor variation, joint factor effects, and changes in model comparisons. Across NHANES, EHRSHOT, and two private health-examination cohorts, covering up to eight model families, features and population generally produce the largest within-source impacts, while missingness has a smaller aggregate impact. Under joint variation, however, features lead sample size in NHANES, whereas sample size leads in EHRSHOT. Small aggregate impacts can also conceal consequential design choices: across the evaluated EHRSHOT missingness configurations, evaluation-only masking produces a task-balanced median AUROC loss of 0.070, compared with 0.002 for all-split masking. In a celiac prediction example, applying 5% missingness within the recent 30-day window reverses the ordering of histogram gradient boosting and a linear model. Dataset design therefore shapes model selection and priorities for data collection; neither follows from sample size or a single benchmark score alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.