acceptodds
Under review as a conference paper at ICLR 2027

Escaping Answer Anchoring: Fixing Shortcut Tokens in Reasoning Models via Context- Replay Gaps

Abstract

Long-form reasoning models are increasingly deployed as autonomous agents, but longer inference and larger context windows do not ensure they use the right evidence. We find that even after extensive reinforcement learning with verifiable rewards (RLVR), reasoning models remain anchored to their earlier answers: when a previous attempt is wrong, they typically restate it rather than produce a different answer. Outcome-level RLVR supplies no direct token-local label, while standard supervised fine-tuning (SFT) weights valid target positions uniformly; neither identifies which tokens depend on the preceding solution. We introduce the context-replay gap, a per-token diagnostic that compares each token’s log-likelihood when the preceding solution is present in context with its log-likelihood when it is absent. Tokens whose likelihood is strongly increased by the replayed solution are separated from those receiving little or negative likelihood change. The gap guides Context-Replay Guided Tuning (CRT) by selecting tokens for an ordinary negative log-likelihood (NLL) loss. This targeted supervision improves mathematical reasoning across model families and benchmarks. Behavioral interventions show that CRT reduces dependence on previously stated answers while preserving sensitivity to the reasoning that supports them. The context-replay gap therefore provides both a diagnostic of answer anchoring and a practical signal for post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.