acceptodds
Under review as a conference paper at ICLR 2027

SFWE: Mitigating Spurious Feature Reliance in LLM Reasoning via Weight Editing

Abstract

Reasoning in large language models can be disrupted by simple perturbations (e.g., adding irrelevant distractors or rephrasing a problem's scenario), suggesting that their reasoning relies on spurious correlations. Identifying and suppressing the spurious features involved in extended chain-of-thought reasoning remains challenging. We introduce Spurious Feature Weight Editing (SFWE), a training-free method that extracts behavioral directions associated with spurious reliance by comparing activations on paired original and perturbed problems, localizes the contributing MLP neurons and attention heads along these directions, and suppresses their influence by removing the direction-aligned components from the selected output weights. Using two small datasets for extracting behavioral vectors and calibration, SFWE improves accuracy on perturbed inputs by 10.9 percentage points on average and accuracy on original inputs by 0.9 points across five open-weight reasoning models and five reasoning benchmarks, reducing the average relative gap between original and perturbed accuracy from 17.4% to 2.2%. Analyses identify two recurring forms of spurious reasoning—over-integrating problem conditions and overthinking—and show that SFWE improves reasoning robustness by suppressing the corresponding behavioral directions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.