Knowing What Will Matter Narrows World-State Coverage in Transformers
Abstract
Accurately predicting one object's state does not reveal how much of an evolving world remains accessible inside a Transformer. We test whether advance knowledge of the evaluated object changes this breadth of accessible state. In a controlled multi-object state-tracking task, paired two-layer Transformers receive identical trajectories and one selected-object final-state label per example. Announced models know the selected identity before the trajectory, whereas Delayed models never receive it during the forward pass. Across the 14 of 20 paired initializations meeting a development-only success criterion in both conditions, selected-state accuracy is closely matched ( vs. ). Yet Delayed models achieve percentage points higher non-target accuracy through native readouts and points higher accuracy with independently fitted linear probes. Counterfactual interventions further show that second-layer attention updates can substantially redirect selected-state predictions while preserving the recipient residual. Development-ranked head sets outperform matched alternatives on held-out interventions, while complementary reversion tests show larger effects on average for the same ranked contributions. Together, these results separate representational breadth from selected-state computation: even among models that answer the evaluated question accurately, knowing what will matter can narrow what remains accessible.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.