When Visual Prompts Rotate: Spectral Guidance for Test-Time In-Context Adaptation
Abstract
Visual in-context learning (VICL) predicts dense outputs from image–annotation demonstrations. Rotating a prompt image and its annotation together preserves the demonstrated task, yet can degrade predictions for an unchanged query. This prompt-only mismatch changes spatial relationships and introduces interpolation, aliasing, and padding; a matched padding control shows that boundary fill alone does not explain the degradation. We propose GIST (Graph-Informed Spectral Tuning), a test-time adaptation method that adds graph-guided attention and error weighting to VICT's support-reconstruction cycle. Spectral Relation Guidance (SRG) uses distances in a low-frequency graph subspace to bias token interactions during both query prediction and support reconstruction. Spectral Adaptation Guidance (SAG) defines a reconstruction metric on a fixed support graph and adjusts its spectral contribution using the preceding residual's frequency composition. The two components guide information exchange and parameter updates without query labels, additional trainable modules, or rotation-specific backbone pretraining. Across 15 corruptions, two prompt settings, and four depth/segmentation benchmarks, GIST reduces average Abs.Rel by 9.9% to 27.3% relative to VICT and increases average mIoU by 0.40 to 1.14 points under matched rotations. Component ablations, attention-bias controls, and geometric diagnostics support these two intervention points while showing that graph structure remains sensitive to real image transformations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.