How Fine-Tuning Data Influences Reliance on Irrelevant Context in Math Reasoning
Abstract
Large language models have been shown to rely on and get distracted by irrelevant context (IC) when solving mathematical reasoning tasks. While there exist many works that show this failure mode, it is not well understood what factors contribute to such reliance on irrelevant context. In this work, we take a first step towards understanding the origins of this failure mode. First, we perform ablations on the irrelevant context and show that models often learn shortcut correlations between artifacts and operations without understanding the requirement of the question, which then causes them to rely on irrelevant context. Next, we discover and show that long chain-of-thought supervision exacerbates reliance on irrelevant context. We also show that introducing irrelevant context during fine-tuning can make these models more robust against irrelevant context. Interestingly, this is true even if the added IC is fully out-of-domain with respect to the original fine-tuning set. We conduct comprehensive generalization tests and also evaluate on more realistic and challenging ICs. Finally, we show that inclusion of such IC within the fine-tuning set does not negatively impact performance on clean test sets, while improving robustness by overcoming reliance on ICs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.