DynaLDI: Dynamic Latent State Estimation with Time-Varying Weak-Signal Reliability
Abstract
Longitudinal language analyses often aggregate heterogeneous weak signals, such as lexicon scores, sentiment measures, and neural classifier outputs, into a single temporal index despite changes in measurement quality and corpus composition. We introduce **DynaLDI**, a dynamic latent-index model that treats these signals as noisy measurements of a shared longitudinal state with autoregressive dynamics, source-specific nuisance effects, and reliability-conditioned time-varying measurement variance. On 480 held-out known-truth simulated longitudinal sequences spanning ten matched and misspecified regimes, the pre-specified DynaLDI configuration reduces latent-state RMSE by 24.0% relative to equal-weight standardized aggregation and by 7.0% relative to an earlier dynamic estimator, outperforming the latter in 91.9% of paired trials. It recovers source loadings with mean cosine similarity 0.999 and time-varying precision with correlation 0.985, while nominal 95% intervals attain 94.9% coverage under matched and near-matched conditions. Applied without retuning to 1.52 million *r/depression* submissions from 2009–2022, DynaLDI reduces several associations with observable corpus drift. Real-data perturbations reveal substantially greater sensitivity to measurement-model re-identification than to conditional state inference. A separate 384-sequence known-truth benchmark shows that the resulting composite sensitivity envelope improves severe-misspecification coverage from 71.6% to 84.4%, although its posterior-variance fraction is not a calibrated local reliability score. Prespecified ecological comparisons show selective temporal convergence with CDC depressive-symptom estimates, NSDUH adult major depressive episodes, and Google Trends searches, while U.S. suicide mortality remains negatively associated. The state-space computational kernel scales approximately linearly with sequence length. The Reddit trajectory is interpreted as a relative latent state of depression-related discourse measurements, not clinical prevalence, diagnosis, or individual risk.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.