BRIDGE: Bilevel Retrieval Credit-Aware Agentic Reinforcement Learning
Abstract
Agentic reinforcement learning (RL) with verifiable reward enhances the capability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM generated tokens and treat external retrieved evidence as environment observations, creating an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than the retriever and motivating the joint training over LLM and retrieval model. In this paper, we show that retrieval and LLM policy learning are order-sensitive, with retriever adaptation followed by policy optimization yielding the largest reward gain. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve this bilevel problem efficiently, we introduce Bilevel Retrieval via Information-Directed Gradient Estimation (BRIDGE), a memory-efficient first-order method motivated by the loss landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving multi-hop accuracy by 9.8% and 3.3%, respectively. It also achieves the best macro-averaged answer accuracy and reasoning quality across five medical QA benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.