Mechanism-Equivalence Routing for Time-Series Foundation Models
Abstract
Sparse mixture-of-experts models scale capacity by conditionally reusing a small subset of parameters, yet the relation that should govern this reuse remains unspecified in time-series foundation models. Routing by trajectories or learned embeddings can merge similar observations with different predictive responses and split a common response across different operating conditions. We formulate mechanism equivalence as a support-conditioned operational reuse criterion: two target-window predictors may share response parameters when their baseline-subtracted multi-horizon responses agree on declared domains under admissible affine coordinate changes. Under canonical squared-error risk and a well-specified equivariant expert family, this relation admits a shared population-optimal response component with a unit-specific horizon baseline. MER operationalizes the relation with finite response probes, a canonical fingerprint, and prototype routing. A controlled CausalTime audit tests routing alignment under designed response families. In frozen target-stage evaluation, MER gives lower mean MSE than pattern routing in all ten directions and than the protocol-matched sparse baseline in all 40 dataset–horizon cells. These end-to-end results support frozen target-stage reuse under the evaluated same-backbone protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.