TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
Abstract
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual condi- tions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception—the foundational bottleneck that anchors multimodal reasoning—risking the degradation of pre-trained reasoning capabilities. We pro- pose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the orig- inal image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B’s MMMU accuracy from 35.79% to 49.32% (+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.