acceptodds
Under review as a conference paper at ICLR 2027

Position Aware Layer Queries for Test Time Training In Vision Language Models

Abstract

Test-Time Training (TTT) adapts models to incoming test samples (e.g. out-of-distribution, OOD) when conventional fine-tuning is infeasible. Existing TTT methods for Vision-Language Models (VLMs) create supervision from several augmented views, each requiring forward (and often backward) passes through the entire VLM, incurring substantial computational cost. We observe that one forward pass with all the intermediate layer outputs already yields far more signal than the final embedding from all augmentations. We introduce **Layer Query Network (LQN)**, a lightweight approach that can adapt a frozen VLM (teacher) in a single forward pass of the VLM via a small model (student). LQN uses **Position-Aware Distillation (PAD)** to mimic the teacher VLM’s intermediate-layer spatial tokens by querying spatial coordinates of intermediate tokens. LQN additionally relies on **Location Consistency Regularization (LCR)**, a self-supervision technique, replacing expensive O(HxW) image augmentation with O(1) coordinate sampling. Integrating these, LQN i) Adapts and improves zero-shot CLIP ViT-B/16 by 9.8% Top-1 on OOD ImageNet, ii) Outperforms the previous best GS-Bias on fine-grained classification by 3.9% Top-1, iii) Faster convergence than TPS for CLIP ResNet-50 (47 mins vs 55 mins) iv) Generalizes adaptation to VLMs like SigLIP, EVA-CLIP, and CoCa, and lightweight students like MLP, ResNet, VGG, and v) Extends to panoptic, instance, and semantic segmentation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.