Where Should LLM Agents Adapt at Test Time? In Context, in Weights, or Both
Abstract
An LLM agent operating repeatedly in the same deployment environment, such as a code repository, a database, or an application suite, can adapt by retaining experience in context, updating its weights, or both. These approaches are usually studied separately, making it difficult to determine their individual contributions and when combining them helps. We compare adaptation through in-context learning (ICL), reinforcement learning with verifiable rewards using GRPO, and self-distillation, individually and in combination, under a common online protocol. Each agent encounters the same task streams, is evaluated on each task before adapting to it, and learns only from its own attempts and environment feedback. We evaluate 14 adaptation configurations and a stateless baseline with Qwen3.5-9B on 109 tasks from four environments drawn from SWE-bench Verified, AppWorld, ScienceWorld, and BIRD-Critic-SQLite, using five task orders per environment. ICL provides the largest standalone gain: its best configuration raises the success rate, averaged across environments, from 28.8% to 49.4%. Adding GRPO to this configuration yields the highest success rate among all configurations at 51.6%. GRPO improves the average success rate under every ICL policy, whereas the effect of self-distillation varies across environments. Further analyses show that GRPO maintains interaction before submission as retained context grows in the database environment, and that adding weight updates to ICL yields larger gains in pass@k at larger k. Finally, GRPO provides larger average gains with an ICL policy that stops admitting new experience once full (ICL-Prefix) than with one that continually replaces the oldest experience (ICL-FIFO). We release a modular framework for evaluating and combining these adaptation mechanisms. Together, our findings provide evidence that context and weight adaptation are complementary and that retaining experience in context does not exhaust the gains available from learning from it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.