Test the Record, Not the Imputation: Model Criticism of Temporal Point Processes from Aggregated Counts
Abstract
Event data such as neural spike trains and activity logs are often recorded only as counts per interval. Exact-time diagnostics for temporal point processes, such as time rescaling, must then impute event times. For history-dependent models, we show that the imputation can decide the verdict: different imputations of the same record can reverse the diagnostic. We instead criticize a model through the distribution it induces on the recorded counts. For continuous-time neural generators this distribution is implicit, so our test, ObsCritic, estimates count forecasts by simulation or particle filtering. It then accumulates signed residuals that show which counts are mispredicted after which histories. Calibrating the whole procedure by Monte Carlo keeps the test valid in finite samples, even when the forecasts are approximate. On synthetic data where frequencies and adjacent transitions of count categories stay fixed, ObsCritic detects changes in higher-order dynamics that imputation-based tests miss. A jointly calibrated library of lag features detects the same alternative. On this matched alternative, signed count residuals are substantially more powerful than a PIT-based diagnostic using the same forecasts. ObsCritic also detects misspecification of a Transformer Hawkes Process and, on real datasets, rejects lower-capacity generators more often. These results support evaluating generators through their distributions on recorded counts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.