acceptodds
Under review as a conference paper at ICLR 2027

GRASP: Graph-Refined Anchored Soft Predictions for Training-Free Vision-Language Model Adaptation Across Stream Structures

Abstract

Training-free test-time adaptation of vision-language models must operate on test streams whose local structure can vary substantially in deployment. Samples observed together may share useful temporal or semantic context, yet heterogeneous batches can make the same relational evidence unreliable. This creates a central challenge for relational adaptation: exploiting local structure without assuming that it is always informative. We demonstrate this failure mode with a simple batch-local graph refinement, which performs strongly on class-correlated streams but degrades sharply under IID ordering. We introduce GRASP (Graph-Refined Anchored Soft Predictions), which retains this local relational structure while controlling its influence. GRASP anchors graph refinement to a cache-refined posterior through a fidelity-regularized objective and attenuates edges whose endpoint posteriors disagree. The graph is restricted to the current batch and discarded after inference; only the bounded cache persists across the test stream. GRASP requires no target labels, gradient updates, prompt optimization, or knowledge of the stream regime. Across 11 image-classification benchmarks and five stream regimes, GRASP achieves the highest cross-regime mean among evaluated methods at . Relative to TDA, which provides its default prediction anchor, GRASP trails by percentage points under IID ordering but improves by points under class-sequential ordering. These results show that local relational evidence can provide substantial test-time gains while limiting degradation under weak local structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.