ER-LVR: RESIDUAL-ADAPTIVE DISPERSION FOR LATENT VISUAL REASONING
Abstract
Latent visual reasoning lets a vision–language model reason over visual evidence without first translating it into words: at designated positions it outputs a continu- ous hidden state in place of a token and feeds it back as the next input embedding. Existing methods train these visual thoughts by regressing each such state onto a target embedding from the model’s own frozen vision encoder. In practice the alignment is imperfect, and the error that remains reflects how well a particular vi- sual thought was formed, yet a single vector carries no trace of it: a visual thought that missed its target re-enters the model’s reasoning looking exactly like one that hit. We show that this is a property of the objective. Squared-error alignment is the degenerate boundary of the energy-score family, at which the score constrains only the mean of the predicted distribution; in the strictly proper interior the scale is constrained as well, and at the optimum it equals the distance from the distribu- tion’s mean to the target, so the spread of a visual thought tracks its own alignment error. We instantiate this as ER-LVR, which replaces the deterministic alignment head with a reparameterized stochastic head trained under an energy objective, and verify that the learned geometry converges to the predicted landmark while two control objectives sharing the same head do not. On spatial planning ER-LVR exceeds the strongest prior latent method by 2.3 points, and by 3.2 with policy optimization of the text policy; it matches or exceeds every latent baseline on out- of-domain visual understanding, and improves joint accuracy on paired semantic queries over a same-backbone baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.