acceptodds
Under review as a conference paper at ICLR 2027

BEYOND THE FINAL ANSWER: CHARACTERIZING AND SELECTING REASONING TRAJECTORIES

Abstract

Sampling multiple reasoning trajectories creates two decisions: which unfinished trajectories deserve more computation, and which terminal answer should be submitted. These stages can benefit from different evidence. A trajectory can merit continuation before its current answer is reliable, while aggregate hidden-state geometry does not determine when substantial updates occur. We show that early answer-support histories predict later correctness even when both the current candidate and an isolated trial answer are wrong. Update timing also helps terminal comparison, including against aggregate-geometric controls. We introduce the Support History Score (SHS), combining candidate dominance frequency, confidence when leading, and historical competitive margin, and the Trajectory Outcome Score (TOS), based primarily on the temporal centroid of representational updates. SHS guides Continuation Selection (CS); two-stage Answer Selection (AS) retains two trajectories with SHS and compares their final submissions with TOS. Across three reasoning models and five mathematical and scientific benchmarks, support history improves continuation selection by 1.4–3.1 percentage points over matched endpoint-only scoring. Within the same retained pairs, TOS improves accuracy by 2.0–4.8 points over uniform choice. Both procedures reduce generated tokens while matching or exceeding majority-vote accuracy in most settings. The reasoning models remain unchanged; supervised two-input TOS extensions use problem-disjoint cross-fitting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.