CopyTestBench: Cross-Channel Retrieval in Transformer Forecasting
Abstract
Real-world forecasting benchmarks can hide important differences between models: similar errors may result from different computations, and failures are hard to interpret when the optimal predictor and irreducible error are unknown. We introduce CopyTestBench (CTBench), a diagnostic, capability-level benchmark with known rules and zero Bayes error, where the correct target always appears in the input window. Rather than only ranking forecasters by aggregate error, CTBench tests whether they retrieve, route, and preserve delayed cross-channel dependencies under pure-copy, daily switching, and rotating switching, with varying distractors and autocorrelation. Ridge regression nearly recovers the exact solution on pure-copy tasks and achieves the lowest overall error, showing that the dependency is observable and linearly solvable. Under the same training protocol, none of the transformer-based forecasters matches this accuracy. Several instead rely on the target level or switching pattern. Some models beat gate-aware persistence under daily switching, but none do so under rotating switching. Controlled experiments identify failures in source routing and normalization. Vanilla Transformer solves one pure-copy dataset under a constant learning rate, while iTransformer remains near the window-mean baseline. Oracle routing, combined with changes to per-channel instance normalization and encoder layer normalization, raises from to , with consistent results across seeds and on an exact-copy task built from real measurements. CTBench therefore connects controlled capability evaluation to mechanistic diagnosis, providing a way to uncover weaknesses hidden by real-world benchmarks and guide the development of stronger Transformer-based forecasters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.