acceptodds
Under review as a conference paper at ICLR 2027

BRIDGE: Bilevel Retrieval Credit-Aware Agentic Reinforcement Learning

Abstract

Agentic reinforcement learning (RL) with verifiable reward enhances the capability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM generated tokens and treat external retrieved evidence as environment observations, creating an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than the retriever and motivating the joint training over LLM and retrieval model. In this paper, we show that retrieval and LLM policy learning are order-sensitive, with retriever adaptation followed by policy optimization yielding the largest reward gain. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve this bilevel problem efficiently, we introduce Bilevel Retrieval via Information-Directed Gradient Estimation (BRIDGE), a memory-efficient first-order method motivated by the loss landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving multi-hop accuracy by 9.8% and 3.3%, respectively. It also achieves the best macro-averaged answer accuracy and reasoning quality across five medical QA benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.