Predictive Factors as Support×Operation Cells: A Strict Compositional Readout for Frozen Encoders under Controlled Interventions
Abstract
Frozen vision encoders encode responses to spatial image edits, yet factor-wise probes do not test whether an operation remains bound to the region where it occurs. They can therefore credit multiple operations through the same predicted slot, a degeneracy we call operation laundering. We introduce an injectively aligned leave-one-cell-out evaluation that requires distinct ground-truth axis values to occupy distinct slots. Our Support–Operation Factorization (SO-OPF) readout separates support salience from operation identity, making the binding between where and what explicit. We distinguish readout capacity under a known factorial assignment from recovery of that assignment using only flat cell labels. Across controlled synthetic scenes and image-disjoint natural images, SO-OPF composes held-out support–operation pairs, while learned assignments retain much of the performance obtained with known structure. Under the same flat-label objective and a shared fixed training budget, SO-OPF outperforms the tested Dense-MLP and Joint-MLP readouts across encoder families and multiple data constructions. Router controls distinguish position-driven from source-driven routing: source features support generalization to unseen layouts, while edit-space transfer characterizes robustness across intervention realizations. These results establish a controlled framework for evaluating readout capacity, assignment recovery, and alignment separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.