Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs
Abstract
LLMs can correct a false premise in isolation yet comply when the same premise is embedded in a routine task. We call this failure correction suppression. We introduce an automated pipeline combining false-premise generation and task-context sampling, with manual verification, to construct paired isolated and contextualized queries. A benchmark of 300 premises spanning seven error categories and 21 domains reveals suppression across eight models, with rates of 19%–90%. Hidden-state and attention analyses on Qwen and LLaMA support an account of knowing but not correcting: task context diverts attention from the false premise as the response trajectory shifts toward compliance in the middle layers. We propose two training-free interventions. Correction Direction Steering (CDS) injects a correction direction at middle layers. Dynamic Payload Amplification (DPA) uses attention changes to locate payload tokens and amplify their representations at the final layer, eliciting correction without injecting a learned semantic direction. Against ITI, DoLa, and TAE on model-specific suppression sets, CDS achieves the highest correction rates across Qwen, LLaMA, and Gemma, with model-dependent capability trade-offs. DPA outperforms these baselines on LLaMA and Gemma and improves MMLU-Pro accuracy over the unmodified model on both models evaluated for general capability. Both methods incur low inference overhead and produce no observed over-correction in the true-premise evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.