TRACE-o1: Reliable Search-Augmented Reasoning via Stage-Aligned Trajectories
Abstract
Search-augmented reasoning enables LLMs to consult external evidence during inference, but search access alone does not guarantee reliable reasoning. Existing interleaved search–reasoning systems often make retrieval and reasoning decisions from local trajectory context, which can lead to fragmented evidence acquisition, accumulated intermediate errors, and unstable final predictions. We propose TRACE-o1, a stage-aligned framework that introduces reliability controls throughout the search-augmented reasoning process. Before reasoning, TRACE-o1 plans an evidence path to guide subsequent retrieval; during reasoning, it audits and repairs the evolving trajectory; and after reasoning, it stabilizes predictions across completed candidate trajectories through consistency-guided selection. Experiments across nine datasets show that TRACE-o1 improves overall answer accuracy and cross-trajectory stability relative to Search-o1-based baselines. Under a matched-K=5 comparison, TRACE-o1 improves macro accuracy from 60.5 to 63.6 over Search-o1 with self-consistency, with a 95% paired-bootstrap confidence interval of [+1.4, +5.0] percentage points. These results support the benefit of combining evidence planning, trajectory repair, and consistency-guided selection within a unified framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.