acceptodds
Under review as a conference paper at ICLR 2027

Learning to Search and Searching to Learn: Joint Knowledge Internalization via Reinforcement Learning

Abstract

Integrating large language models (LLMs) with external search engines has significantly enhanced their ability to solve knowledge-intensive tasks. While search-integrated LLMs optimized via reinforcement learning (RL) excel at learning when and how to retrieve information, they typically treat retrieved text strictly as read-only context. To ensure training stability, these retrieved tokens are often masked from the policy-gradient objective, which inherently prevents models from internalizing the factual content. However, directly unmasking these tokens introduces severe noise and destabilizes the early stages of optimization. To resolve this masking dilemma, we propose LfSearch, an end-to-end RL framework that shifts the paradigm from merely learning to search toward searching to learn. LfSearch jointly optimizes search behavior and knowledge internalization through a unified dual-objective loss. By introducing an adaptive masking curriculum, the model is permitted to absorb external facts only after mastering reliable tool use. Furthermore, we apply asymmetric advantage normalization to ensure factual content is strictly reinforced rather than penalized. Across seven single-hop and multi-hop QA benchmarks, LfSearch successfully circumvents the severe performance degradation typically caused by naive unmasking. Empowered by this internalized knowledge, it consistently outperforms baselines like Search-R1, demonstrating that learned facts directly translate to superior dynamic search performance. Finally, evaluations on knowledge-retention sets explicitly confirm that LfSearch effectively distills retrieved facts into parametric memory, enabling the model to genuinely learn from its environment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.