From Partial Likelihood to Pairwise Preference: Deep Survival Analysis with Bradley–Terry Models
Abstract
Deep learning has become central to modern survival analysis, with applications spanning clinical trials, operations research, and economics. A defining—and largely underexplored—feature of many such applications, most acutely in medical imaging, is that the covariate dimension far exceeds the number of available subjects (): a single CT volume already contains on the order of voxels, whereas a cohort rarely exceeds a few thousand patients. This regime forces very small training batches, which interact poorly with the Cox proportional hazards model: its partial likelihood couples all subjects through a shared risk-set denominator, making minibatch gradients biased and stochastic optimization unreliable exactly when batches are smallest. We make a simple observation with broad consequences: fitting a Bradley–Terry model in place of the Cox partial likelihood recasts survival training as a sum over independent pairwise preferences, yielding unbiased minibatch gradients while targeting the same risk ordering. Casting survival analysis as preference learning further lets us import the extensive toolbox developed around the Bradley–Terry model in RLHF—graded (fine-grained) labels that exploit the magnitude of survival-time gaps. On a 3D medical imaging dataset, our method matches or surpasses Cox-based deep survival baselines, with the largest gains in the small-batch, high-dimensional regime where partial-likelihood training degrades most.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.