When Sparse Routing Looks Additive: Hidden Pairwise Structure Changes Route Choices
Abstract
An additive fit to mean sparse-route loss need not preserve token-local effects or valid route choices. We audit temporal preservation and choice validity separately using complete four-band intervention surfaces and exact functional ANOVA. Beyond guaranteed energy attenuation under averaging, pairwise effects retain less normalized energy than unary effects on average. In a fresh 40-source, two-model study, the mean within-surface pairwise share of nonadditive gold-NLL energy is 86.72%; retaining pairwise terms reduces pooled mean hindsight regret versus unary by 77.5%, equally weighting separately evaluated NLL/KL endpoints. On 160 further documents, degree-two projection fitted to the same 70 feasible native Quest routes as unary reduces mean native-reference KL hindsight regret by 59.9% under fixed Llama/H100 execution. A prespecified support-construction intervention on 80 new documents narrows the retention gap at fixed K and mean overlap. With per-document calibration, two separate 48-document eight-band studies confirm held-configuration prediction at cohort-mean and document levels, including 28.69% lower mean document NLL RMSE versus unary in the prespecified Llama-interleaved condition. Better prediction can still yield a worse route. These findings motivate evaluating approximation-induced choices directly and reporting support construction when comparing preservation gaps.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.