A Dual Scalar Field on Manifold for Recursive H-Neurons Detection and Mitigation in Large Language Models
Abstract
Hallucination-Associated Neurons (H-Neurons) refer to a pro-hallucinatory subset of separable neurons. Current methods identify this subset within the activation space by deploying linear probes to impose external statistical boundaries. Such an external empirical estimator lacks geometric interpretability, fundamentally struggling to yield a unified hallucination mitigation framework, while the exhaustive fitting of probes across successive layers incurs prohibitive cross-layer computational costs. Accordingly, we dually derive a scalar field on the manifold and recursively propagate it through the network layers, with its geometric properties intrinsically dictating H-Neurons detection and the mitigation direction. Specifically, we theoretically prove the equivalence between the tangent vector of the scalar field and that of the geodesic on the support connecting hallucinated states to truthful states. This fundamental equivalence enables both the identification of H-Neurons and the comparison of hallucination degrees: (1) decomposing the tangent vector onto activation axes allows component polarities to identify H-Neurons by revealing their pro-hallucinatory dynamical roles and (2) differentiating along the geodesic yields a non-negative derivative, establishing that the scalar field is monotonically non-decreasing along the geodesic trajectory toward truthful states. Subsequently, guided by tangential equivalence and monotonicity of the scalar field, we formalize hallucination mitigation as a path integral process governed by a first-order ordinary differential equation. This motivates a local trajectory along which hallucinated activations follow the ascent direction of the scalar field, thereby steering them toward regions associated with truthfulness. Finally, to alleviate cross-layer computational overhead, we introduce a manifold pullback mechanism to recursively derive high-quality scalar field candidates with a bounded approximation error, thereby reducing the need for exhaustive layer-wise fitting. Experiments demonstrate that our framework can accurately detect H-Neurons and significantly mitigate hallucinations in large language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.