acceptodds
Under review as a conference paper at ICLR 2027

Bring the Answer into View: Answer-Bearing Frontier Training for Search Agents

Abstract

Outcome reinforcement learning directly optimizes final-answer correctness, but its terminal signal conflates a trajectory that never acquires the needed information with one that retrieves it but fails to use it. Fine-grained rewards offer denser credit, yet a search trajectory mixes exploration, acquisition, verification, and stopping, for which one proxy need not have a consistent meaning. We study a complementary principle: apply a simple, clear reward at a position where its intended consequence is directly observable. The Answer-Bearing Frontier (ABF) is the state immediately before a tool observation first contains the reference answer; the realized transition witnesses that answer-bearing information is reachable in one step. Answer-Bearing Frontier Training restores these states, samples alternative one-decision continuations, and rewards immediate acquisition, while full-trajectory Outcome updates optimize reaching useful states, integrating information, stopping, and answering. With local supervision only at selected turns, ABF Training achieves 37.4% and 48.8% average exact match across six benchmarks at 2B and 9B, with margins of +2.3 and +1.7 points over the strongest comparators and leads 9 of 12 dataset–scale settings. Relative to Search-R1 at 2B, answer-bearing coverage rises from 53.8% to 61.8%, while exact match given such information changes from 51.9% to 53.0%. The gains extend beyond the answer types used for local training. These results show that a strong training position can turn a narrow local reward into stable end-to-end gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.