SUTA: Towards Online Adaptation for Self-Improving Large Language Models
Abstract
Deployed language models are typically static: experience accumulated while processing one input does not improve their behavior on later inputs. Test-time adaptation offers a practical form of online self-improvement by allowing a model to update persistent parameters from unlabeled deployment data. For autoregressive LLMs, however, treating every token in a self-generated trajectory as an equally reliable update signal can reinforce confirmation bias and cause reasoning collapse. We introduce **S**elective **U**ncertainty-aware **T**est-time **A**daptation (SUTA), a lightweight parametric self-improvement mechanism for autoregressive LLMs. SUTA uses predictive entropy to identify localized uncertainty bottlenecks in its own generations and restricts cumulative adapter updates to these high-entropy positions. An information-gain contrastive objective encourages the updated model to reduce uncertainty through the input context rather than context-free shortcuts, while a frozen-prior preservation term constrains drift during online updates. Across mathematical reasoning, domain knowledge, and coding benchmarks, SUTA consistently outperforms representative test-time optimization baselines on pre-trained and instruction-tuned models. Its gains grow with model scale and remain complementary to inference-time sampling and voting. These results show that selectively resolving uncertainty in self-generated trajectories is an effective and stable route to test-time self-improvement under distribution shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.