acceptodds
Under review as a conference paper at ICLR 2027

Naturalness Is in the Heads: Localizing and Enhancing Translation Naturalness in LLMs

Abstract

Large language models can generally follow translation instructions; the central challenge lies in producing translations that read naturally in the target language. To inform quality improvement, mechanistic analysis should explain how models choose among correct translations, not only how they follow translation instructions. Naturalness unfolds over complete sequences and admits multiple valid wordings, so neither a single-token readout nor a single reference adequately captures the behavior. We find an exceptionally sparse causal circuit: intervening on only the top-ranked 1% of attention heads shifts translation naturalness, while steering the localized circuit enhances it. We localize these heads using the equivalence-class score difference (Equi-DIFF), which contrasts cardinality-normalized scores over classes of valid multi-token translations and reduces exactly to logit difference in the single-token singleton case. Comparisons against word-level translation and sequence-singleton baselines show that complete-sequence scoring materially changes head rankings, while multi-completion classes improve calibration and reproducibility. Held-out knockout reduces natural-style preference before broad quality degradation, whereas activation steering enhances judged naturalness (69–87% model-level decisive win rate). Across three models, 12 translation directions, and two corpora, this effect recurs in most settings, providing consistent causal evidence for naturalness heads despite varying adequacy trade-offs. Code will be made available after publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.