Prior Directions: Separating Geometric Identification from Behavioral Efficacy in Visual Revision
Abstract
Vision-language models can remain anchored to stale verbal priors even when current visual evidence supports a different answer. We use this controlled revision failure to ask whether the subspace that best captures recurrent prior-induced representation change is also uniquely more behaviorally effective. Across four models, paired prior and reference states define Prior Directions, a compact low-rank geometry that, at rank 32, spans only 0.78% of the state dimensions yet captures 49.7% to 64.0% of held-out displacement energy. Matched correspondence improves held-out capture at every tested rank in all four models; across 20 permutation seeds, this advantage remains positive in all 15 available model–rank settings. Yet the geometric advantage does not translate into a comparable intervention advantage: matched and Pairing-Permuted edits restore 60/66 and 59/66 clean prior-induced failures, respectively. We trace this discrepancy to the geometry of the realized edits. Pairing-Permuted edits remain strongly aligned with the matched Prior Directions; their aligned component restores 59/66 cases, whereas an equally norm-matched orthogonal residual restores only 20/66. These results show that the construction that best identifies recurrent representation change need not be uniquely more behaviorally effective, distinguishing global subspace identification from the geometry actually used by an intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.