acceptodds
Under review as a conference paper at ICLR 2027

Censoring as an Outcome: One Diffusion for Survival Prediction and Cohort Synthesis

Abstract

Survival models estimate when the event happens and treat when the patient left as a nuisance. That discard is why none of them yields a synthetic cohort a statistician can analyse: such a cohort needs the dropout mechanism too, so the target is the law of both latent times. Censoring at random factorises that law and makes it symmetric, an event record right-censoring the censoring time exactly as a censored record right-censors the event time. The censoring channel is then a fitted distribution rather than a weight, and follow-up becomes a quantity one can set instead of a property of the dataset: the model resamples a cohort under a dropout regime it never saw, which is what a power calculation or a planned extension asks for. A model of the observed pair has one lever, truncation, and truncation only shortens. We fit the law with a two-channel score-based diffusion on , trained by variational EM whose E-step draws the unobserved channel from a truncated normal in closed form; the objective upper-bounds the censored-channel negative log-likelihood, and the proposal's slack and the chain's consistency are bounded with it. On six real cohorts the same model leads nine baselines on macro integrated Brier score, time-dependent AUC and Brier resolution, four comparisons resolving at the cohort level, and composed with a tabular backbone it is best on every downstream-utility median we score and beats four competing generators in all six cohorts. Rescaling the fitted dropout curve leaves the event curve inside the estimator's own noise, where truncation matched to the same censoring rate moves it ten times further. Dropping the channel costs IBS, concordance, Brier resolution and reliability in five of six cohorts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.