acceptodds
Under review as a conference paper at ICLR 2027

Vernier: Diagnosing and Mitigating Lexical Sensitivity in Causal Reasoning

Abstract

Lexical renaming exposes a gap between specifying a causal problem and eliciting a stable answer from a language model. We study this gap through Vernier, a paired-view LoRA update that combines supervision on original and placeholder prompts with answer-token consistency. The analysis separates three outcomes: improved task performance, agreement across lexical views, and the representation measurements used to explain them. Recorded multi-model evaluations show substantial gains in both views for several Qwen and Llama models, with stronger transfer on CRASS than on e-CARE. An independently traceable 900-question evaluation confirms large paired-view accuracy gains, while also showing that such gains need not reduce the absolute lexical gap. Probes measure the transfer of lexical-content decoders, and donor-controlled patching demonstrates strong control of answer identity at an intermediate layer. Together, the results locate a useful target for adaptation at the answer interface and establish why task improvement, lexical consistency, and mechanistic explanation must be evaluated separately.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.