acceptodds
Under review as a conference paper at ICLR 2027

ASCT: Agentic Search Critical Training

Abstract

Reinforcement learning (RL) holds great promise for training autonomous agents in sequential decision-making, yet outcome-based methods face a fundamental credit assignment problem in long-horizon interactions: sparse rewards provide no signal about which intermediate actions were effective. Existing remedies are costly. Process Reward Models (PRMs) require training and running an auxiliary model to score every intermediate step at inference time, incurring substantial runtime overhead. Reward shaping relies on hand-engineered multi-objective functions, and value-based credit assignment suffers from unstable value estimation over long horizons. We propose Agentic Search Critical Training (ASCT), a two-stage RL framework built on a simple insight: teaching an agent to critically judge which candidate action is better and to reason about why it is better yields a higher performance ceiling than teaching it to imitate correct answers. Rather than scoring steps with an auxiliary model at runtime, ASCT decomposes multi-step search trajectories into three hierarchical perception levels—environment cognition, strategic decision-making, and tactical execution—and constructs multiplechoice discriminative tasks at each level via trajectory truncation. In Stage 1, hierarchical rewards inject these critical priors in a controlled setting; in Stage 2, all scaffolding is removed and the agent explores autonomously with only outcome rewards, leveraging its internalized analytical capabilities. Experiments on multi-hop QA and DeepResearch benchmarks show that ASCT consistently outperforms GRPO and other strong baselines across multiple model scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.