acceptodds
Under review as a conference paper at ICLR 2027

Test-Time Adaptation under Resource-Constrained Deployment via Attention-Head Alignment

Abstract

Distribution shifts can degrade the response quality of large language models (LLMs), while resource-constrained deployments may lack relevant corpora or the memory needed to host additional pretrained evaluators. Many existing adaptation methods also assume access to multiple test samples from a coherent target distribution to stabilize updates, an assumption that may not hold for queries from different users or domains. We propose per-sample test-time adaptation through attention-head alignment. Brief prefix exploration exposes higher-quality candidates, while prompt–candidate attention-head activation differences provide transferable signals for selecting them. A lightweight scorer calibrated offline on a small set of generic prompt–reference pairs uses these signals to select a response for each query. Both the LLM and scorer remain fixed at deployment, enabling forward-only adaptation without retrieval, additional pretrained evaluators, or information retained from previous queries. Across three LLMs and six benchmarks, our method achieves higher mean ROUGE-LSUM than greedy decoding and the evaluated episodic update baselines for the evaluated backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.