From Reflection to Discovery: Active Knowledge Acquisition for Training Context
Abstract
Most existing large language models (LLMs) are expensive and time-consuming to adapt after deployment, especially when a task requires newly produced information or niche domain knowledge. Recent work has shown that LLMs can iteratively update their own context to adapt to downstream tasks without updating their weights. However, relying solely on their intrinsic knowledge limits their ability to acquire missing task-relevant information: feedback may identify errors without providing the knowledge needed to correct them. In this paper, we equip these LLMs with Wikipedia search and browser tools to actively acquire missing information during context updates. Simply adding these tools to standard sequential context training can degrade performance. We therefore adopt a beam-search-style procedure that explores multiple candidate contexts and retains states selected by validation feedback, including the previous best context. We evaluate this combination on low-resource translation (FLORES+), healthcare (HealthBench), and reasoning tasks (LiveCodeBench and Humanity's Last Exam), observing gains over the evaluated non-information-seeking baselines. Further analyses examine data efficiency and hyperparameter sensitivity on low-resource translation, and show transfer of the optimized context from Gemini-2.5-Flash to Gemini-3-Flash.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.