acceptodds
Under review as a conference paper at ICLR 2027

The Hidden Optimizer in Spectral Time-Series Forecasting

Abstract

Multi-branch architectures in spectral time-series forecasting are commonly presumed to expand representational capacity. This paper demonstrates that such structural additions often merely counteract a hidden optimization pathology rather than enrich expressive power. To isolate this phenomenon, we utilize a modular Phase-Structured Predictor (PSP) that decouples phase, frequency, and input-conditioned pathways, establishing a controlled diagnostic testbed. Through this architecture, we reveal that standard adaptive optimizers (e.g., AdamW) leave an acute update-scale attenuation in forecast space due to the inverse Fourier transform, severely throttling terminal kernel adaptation and corrupting upstream representations. We prove that under whole-model closure, multi-branch dynamics can be rigorously collapsed into a single effective kernel. Grounded in this equivalence, we introduce SpecScale (Spectral Scale Optimizer): theoretically, its Exact formulation strictly reproduces multi-branch training dynamics step-for-step with zero trajectory divergence; practically, it yields an efficient, branch-free rule that applies an analytic frequency multiplier directly to base-kernel AdamW updates. On a four-cell diagnostic panel under closure, SpecScale recovers of the multi-branch empirical gain. On the full 32-cell weighted benchmark, SpecScale reduces test MSE by (improving 27 of 32 settings under equal-budget selection) while matching a 15-setting hyperparameter grid search with fewer trials and zero inference overhead. Furthermore, readout corrections consistently transfer to three prominent external backbones, confirming that readout scale attenuation is an architecture-agnostic optimization bottleneck.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.