Diagnosing and Repairing Long-Horizon Failure in Transformer Queueing Simulators
Abstract
Neural event simulators are evaluated on short test histories but deployed over much longer rollouts, where accurate one-step predictions need not ensure faithful trajectories. Standard test loss leaves predictions beyond the training length unexamined. We investigate whether predictions on shared long reference histories can diagnose these failures and guide corrective supervision. In an M/M/1 queue, we evaluate transformers trained on -event sequences using rollouts of events per model, and compute deep-prefix drift (DPD), the mean predicted queue increment beyond the training length, on simulator-generated reference histories. Across failing checkpoints and three fine-tuning seeds, we compare deep supervision with matched shallow supervision and continued training, and on three of them compare exact conditional targets with fresh simulator draws and recorded continuations. Despite near-oracle short-history test loss, models cross a fixed queue boundary on – of rollouts, with of first crossings occurring beyond the training length. On evaluation models, DPD ranks crossing counts at Spearman and flags of severe failures with false positive. DPD requires neither candidate-model rollouts nor exact conditional probabilities, using model predictions on shared simulator-generated reference histories. Deep supervision produces fewer crossings than both controls in all model-by-seed comparisons, by at least of rollouts, and fresh simulator targets outperform both controls in all tested comparisons, while recorded continuations leave at least more crossings than exact targets. These findings support shared long-history evaluation as a parallel diagnostic for learned event simulators and show that supervision beyond the training length can reduce failures hidden by short-history test loss.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.