acceptodds
Under review as a conference paper at ICLR 2027

Perplexity-Guided Trajectory Repair Shows No Advantage over Length-Matched Controls

Abstract

Recovery from a failed agent trajectory can restart the task or keep a prefix and regenerate its suffix. We test whether stored-token perplexity helps choose where to regenerate, beyond positional controls, in ReAct-style QA agents. In a pre- specified 250-question HotpotQA study, two-step backtracking from the highest- perplexity step achieves 9.73% repair success versus 11.21% for restart (95% in- terval −4.13 to +0.88 points). A diagnosis/replay follow-up, a frozen two-model, three-dataset replication and a deadline study find no clear advantage in any pre- specified Holm family. A length-matched swap, which gives each failed question the uncertainty-chosen origin of an equal-length partner, holds origin distributions exactly equal by construction and is also null, including two confirmatory cells with enough discordant pairs to resolve a large effect: −0.17 points (p = 1.000) at 32B and +1.48 (p = 0.397) at 72B. All experiments pass replay, scoring and allowance audits. These results limit claims for this policy in offline QA; they establish neither equivalence nor an advantage under equal total compute.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.