Hyperspherical Steering with Optimal Transport Weighted Dual Prototypes for LLMs
Abstract
Training-free hyperspherical activation steering enables lightweight inference correction for large language models (LLMs) via norm-preserving rotation, yet its geometric framework has inherent limitations: uniformly aggregated centroids are vulnerable to boundary noise; the single-difference prototype implies an antipodal constraint that causes systematic bias in rotation direction; global parameters fail to balance the demands of discriminative and generative tasks. This paper proposes an optimal transport-weighted dual-prototype hyperspherical steering method. In the offline pre-computation stage, we first construct prototypes via hyperspherical optimal transport weighting to suppress boundary noise and enhance inter-class geodesic separation. We then build an independent dual-prototype paradigm to lift the enforced antipodal constraint, using truthful prototype as correction anchor to eliminate rotation direction bias at its geometric root. During online inference, a task-aware von Mises-Fisher gating mechanism is adopted to differentially regulate intervention strength. Experiments on LLaMA-3.1-8B and Qwen2.5-7B across six benchmark datasets show that the proposed method introduces negligible extra computational overhead, and achieves average improvements of 5.33% and 10.14% over SOTA baselines on multiple-choice and open-ended generation tasks respectively. Rank collapse analysis verifies that our method induces milder representation collapse and preserves the LLM’s native representation space more completely. This work improves the theoretical framework of hyperspherical steering from its geometric fundamentals, providing a more self-consistent technical path for training-free lightweight hallucination correction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.