acceptodds
Under review as a conference paper at ICLR 2027

Loss-Aware Test-Time Subspace Adaptation for Reduced-Rank LLMs

Abstract

Low-rank compression reduces deployment memory but fixes the model subspace before the request is known. We ask whether a deployed reduced-rank language model can adapt its stored factors at inference time without reconstructing or retaining dense weights. We introduce request-local loss-aware test-time adaptation (TTA) in the deployed low-rank subspace: observed prompt tokens define a masked causal loss, request gradients choose update support, and a hard Kullback–Leibler (KL) trust region limits functional drift. Across Qwen3 models, test-time adaptation consistently improves held-out WikiText likelihood and improves GSM8K in selected gauge/support settings, while ARC-Challenge chat remains statistically unchanged despite substantial prompt-loss reduction. Matched-KL ablations further show that gauge, factor side, and block support can change downstream behavior even when functional displacement is nearly fixed. Together, these results identify update geometry and surrogate–task alignment as central design variables for adapting compressed models at test time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.