S2R: Selective-Stability Repair in Frozen Vision-Language Models
Abstract
Vision–language models can be unstable when an answer should stay the same and insensitive when its visual support is weakened. We call the desired behavior *selective stability*: agreement across answer-preserving text and image changes, coupled with reduced confidence when proposed visual evidence is attenuated. **S2R** turns this criterion into a test-time repair signal. It reads answers at several points along the base reasoning trajectory and fits a temporary, hard-sparse edit of visual-token states in an otherwise frozen model. Hyperparameters and gate thresholds are selected on separate calibration data. On Qwen3-VL-8B-Thinking, S2R raises macro-average accuracy from 72.4% to 75.5%, while S2R-Full reaches 75.7%. The improvement transfers to two other frozen 8B backbones, and ablations show that preserving views, the directional evidence constraint, and localized support contribute complementary gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.