CALIPER: Compute-Aware Local Intervention and Targeted Policy Updates for Tool-Augmented Agents
Abstract
Agentic RL commonly samples, compares, and updates full trajectories. Under sparse terminal rewards, more rollouts need not yield more useful supervision: full-trajectory generation consumes the budget, one outcome conflates multiple behavioral changes, and trajectory-level advantages may update decisions unrelated to the observed outcome change. This problem is pronounced in tool-integrated reasoning (TIR), where local changes arise both from autonomous decisions and from responses to newly received tool feedback. The challenge is to choose where to spend computation, which local changes merit comparison, and how to turn their evidence into localized training credit. We propose CALIPER. Under a fixed rollout budget, CALIPER first allocates a portion to full-trajectory exploration, with resulting trajectories serving as training samples and online probes of the current policy. From failed trajectories, CALIPER identifies two types of candidate positions, autonomous model decisions and tool-feedback-driven responses, and estimates their per-compute value from historical success rates, uncertainty, and generation cost. It then adaptively allocates the remaining budget between full exploration and local validation. For selected positions, CALIPER preserves the preceding interaction history and tool state while resampling only subsequent behaviors, yielding strictly verifiable same-state comparisons. During training, CALIPER combines global trajectory advantages with reliability-calibrated local Leave-One-Out (LOO) credit, restricts local credit to the regions where behavior actually changes, and uses negative local evidence to selectively strengthen policy updates. Even when all full trajectories fail, local comparisons can provide non-zero and spatially localized training signals. Across tool-augmented reasoning benchmarks, CALIPER overall outperforms GRPO and other strong baselines, demonstrating improved training-information efficiency under limited rollout budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.