When the Rationale Disagrees with the Log: Diagnosing and Repairing Chain-of-Thought for Sequential Recommendation
Abstract
Chain-of-thought reliably helps in mathematics, code and logic yet is repeatedly reported to be neutral or harmful in recommendation. We show that which of the two occurs is decided by a single quantity measurable before any reasoning supervision is trained, and we derive it in closed form. Because a decoder-only model emits the rationale and the answer through the same readout, the reasoning and ranking gradients satisfy an exact identity whose sign is governed by one indicator: whether the item the rationale argues for is the item that was logged. In mathematics the indicator equals one by construction and the two gradients reinforce. In non-verifiable, multi-modal prediction, where many continuations are legitimate but only one is recorded, it collapses to the overlap between the true continuation law and the teacher's; the gradients then provably conflict. Three further measurements confirm the mechanism: the reasoning-ranking cosine turns negative exactly when predicted, deepens toward the layers at which the model commits to an answer, and leaves reasoning states nearly answer-agnostic under a linear probe. The identity is also constructive, naming token-level rationale imitation as the sole negative driver and predicting that removing it reverses the sign. It does: conflict-aware training (Care) and outcome-reward reasoning restore answer decodability and turn chain-of-thought from a penalty into a gain over direct supervised fine-tuning. A corollary for practice is that the prevailing remedy, replacing discrete traces with latent ones, relocates the imitation target rather than removing it, which is why a change confined to the loss outperforms it. Reasoning is not the wrong tool for recommendation; imitating rationales is.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.