acceptodds
Under review as a conference paper at ICLR 2027

A Smooth Temporal Low-Rank Basis for Neural Estimation of Cumulative Longitudinal Means

Abstract

To characterize how longitudinal outcomes are driven by the covariate history, we consider modeling the conditional mean of the outcome at visit , , with nonlinear and unknown functions of the concurrent covariate vector , . When is additive elementwise across , the estimation is well studied in the setting of distributed-lag non-linear models (DLNM) at the standard spline rates. When the structural assumption of additivity on is not imposed, little work addresses the sample sizes typical of cohort studies, hundreds to a few thousand subjects. In particular, DLNM cannot capture interactions, which costs prediction accuracy that no increase in sample size recovers. On the other hand, sequence models, such as recurrent networks, Transformers and state-space models, specify an unrestricted function of the covariate history and impose no structure, but must learn the functional form from the data; and the direct neural encoding of the cumulative mean model, with a separate network for each visit, has an estimation variance that grows with the number of visits. We propose SmoothCum, a neural estimator that keeps the cumulative form of the model, shares one feature-extracting network across all visits, and expresses the visit-specific output layer through B-spline basis functions of time. For a sample-split training procedure, we establish an exact decomposition of the mean squared error into squared bias and variance, where the variance term grows with the number of basis functions and the width of the network's last hidden layer. Simulation experiments use this decomposition to quantify the bias-variance trade-off of SmoothCum and to compare it with competing methods. The measured variance of SmoothCum is linear in the number of basis functions, with slopes scaling correctly in sample size and signal-to-noise ratio; the cross-validated number of basis functions adapts to the smoothness of the true time-varying coefficient; the advantage over the estimator with a separate network per visit grows with the number of visits, by a factor of to ; and under irregular observation times, where a separate network per visit cannot be built, SmoothCum outperforms all competing methods, including DLNM, by a factor of to . An analysis of cohort data agrees with the bias-variance decomposition: the advantage of SmoothCum appears at cohort sample sizes and disappears in cohorts of thousands of subjects, where the variance it removes is small. When the model class is unknown, combining DLNM and SmoothCum by model averaging, with weights chosen by cross-validation over subject-level folds, was not worse than a single model in any setting we evaluated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.