acceptodds
Under review as a conference paper at ICLR 2027

What Does Answer-Based Co-Training Teach LLM Search Agents?

Abstract

Agentic search enables large language models to solve knowledge-intensive tasks through repeated retrieval and reasoning. Reinforcement learning trains these agents with final-answer rewards, but sparse feedback gives limited guidance for intermediate search decisions. Answer-based co-training addresses this problem by training a model to answer at intermediate steps and using improvements in answer quality as process rewards. However, how to combine these two training signals into a stronger search agent remains unclear. We investigate four design questions through controlled comparisons with full co-training. First, for parameter sharing, we compare shared and separate models: sharing parameters improves the deployed policy's answers on the same evidence by applying intermediate-answer updates to the policy itself. Second, for training signals, we separate intermediate-answer training from process rewards: intermediate-answer training improves early responses and robustness to changed answering instructions. Third, for reward placement, we compare stepwise delivery with terminal aggregation: accumulating process rewards at the trajectory end provides an effective alternative while preserving their sum. Fourth, for feedback information, we compare agent performance after training with or without prior reasoning in the answer model's input. We separately compare scoring criteria: continued-search scoring identifies useful clues better than immediate-answer scoring by testing whether the policy can follow them to the answer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.