SPECTRA: Conditional Prediction with Side-Input-Invariant Structural States
Abstract
Predicting human chess moves calls for adapting to a player's strength. Can those predictions change while the model's board computation stays fixed? We study this question with SPECTRA. At each layer, board tokens read only board tokens; the rating-dependent side token reads the board tokens and itself; the prediction token reads all three types. Under the stated input and layer assumptions, changing the rating cannot alter any board state at fixed parameters. On the same 19,546 games, three seeds place SPECTRA top-1 percentage points above late fusion and points below unrestricted attention (mean seed SD). These comparisons evaluate complete architectures, not an isolated connection. Reusing board states makes 26 rating queries faster than ordinary SPECTRA execution at batch 64, while late fusion remains faster in absolute terms. Other input partitions preserve the same invariant but have different predictive costs, including a substantial loss on Waterbirds. Keeping board states fixed permits reuse, but does not by itself improve prediction; the measured tradeoffs depend on the task and input partition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.