Layer-Wise Semantic Evidence Rewriting under Unified Mask-Token Training
Abstract
Extending vision-language models to predict masks can also change their answers to semantic questions. What changes in the evidence behind those answers? We study layer-wise semantic evidence rewriting between reference and unified checkpoints. Our framework decomposes a decision margin into input and residual-block contributions. The decomposition accounts for normalization and is exact for a fixed forward pass. Comparing representations under checkpoint-sourced readouts separates representation-side changes from readout-side changes. A calibrated linear transport then maps block evidence between checkpoints. Across five model-benchmark comparisons spanning two backbone families, representation-side drift exceeds readout-side drift. Full transport reduces reconstruction error by 30.6-69.0% relative to uniform rescaling on the original evaluation splits. A retrospective group-disjoint reanalysis finds positive sample-level associations in four settings. Full transport also yields lower observed mean absolute errors than scalar, diagonal, and probe-mean baselines. Its advantage over diagonal transport persists when both include layer-wise intercepts. In HiMTok, selected late-layer groups are sensitive to residual-write rescaling, but no more sensitive than the last three blocks. These findings reveal a layer-wise structure in semantic change that aggregate accuracy alone cannot show.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.