Vernier: Diagnosing and Mitigating Lexical Sensitivity in Causal Reasoning
Abstract
Lexical renaming exposes a gap between specifying a causal problem and eliciting a stable answer from a language model. We study this gap through Vernier, a paired-view LoRA update that combines supervision on original and placeholder prompts with answer-token consistency. The analysis separates three outcomes: improved task performance, agreement across lexical views, and the representation measurements used to explain them. Recorded multi-model evaluations show substantial gains in both views for several Qwen and Llama models, with stronger transfer on CRASS than on e-CARE. An independently traceable 900-question evaluation confirms large paired-view accuracy gains, while also showing that such gains need not reduce the absolute lexical gap. Probes measure the transfer of lexical-content decoders, and donor-controlled patching demonstrates strong control of answer identity at an intermediate layer. Together, the results locate a useful target for adaptation at the answer interface and establish why task improvement, lexical consistency, and mechanistic explanation must be evaluated separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.