From Units to Layers: Structural Intervention for LLM Hallucination
Abstract
Large language models (LLMs) are known to exhibit hallucination, producing fluent but factually incorrect outputs. Existing approaches largely treat hallucination as a black-box phenomenon, offering limited internal insight. Prior mechanistic studies of hallucinations have focused on individual computational units, whose ablation tends to suppress the targeted behavior at the cost of performance on unrelated tasks. We argue that this cost is structural. Individual units may participate in many functions at once, which removal almost necessarily perturbs unrelated computation. In this work, we propose a structure-level alternative. We localize hallucination-related units via attribution and causal validation, and map them onto the model's layer-wise structure. It reveals that hallucination concentrates in specific structural regions that are stable across tasks. Building on this insight, we introduce a set of structure-level intervention operators, including Freeze, Swap and Insert, to modify layer-wise computation while preserving the overall model structure. To our knowledge, this is the first systematic class of structure-level operations for hallucination mitigation in LLMs. We systematically evaluate these operators against unit-level ablation in both intra-task and cross-task settings, spanning benchmarks for knowledge-based question answering, grammatical acceptability, and reasoning. Our solution achieves comparable to or stronger performance than the state-of-the-art, indicating that hallucination is governed by higher-level computational organization rather than by isolated units. It highlights structure-level intervention is a more principled basis for behavioral control in LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.