acceptodds
Under review as a conference paper at ICLR 2027

When Attention Heads Herd: Spectral Release for Faithful Multimodal Reasoning

Abstract

Long-form multimodal reasoning improves capability but can let a fluent chain drift from visual evidence. Existing test-time remedies target output distributions, fixed heads, or attention intermediates that stock fused scaled-dot-product attention (SDPA) does not expose. We identify a distinct population-level spectral pattern in residual-stream head contributions. Under spectral herding, temporal innovations remain concentrated around a dominant residual-space mode even as its carrier heads change. The leading spectral share persists through carrier turnover, exceeding a row-energy-preserving permutation null across three models. Yet spectral peaks occur in both factual and hallucinated traces: factual events release them as innovation rises, whereas hallucinated events retain them, producing a failed spectral release. We introduce SPLIT (Spectral Peak Liberation via Instantaneous Turnover), a training-free, single-pass controller that tracks this mode and redistributes gain toward the head bulk with exact reference-energy conservation. We characterize conditions for sketch stability and temporal tracking, proving permutation-equivariant steepest release. Across five models and eight benchmarks, one shared setting improves the underlying models and outperforms five recent baselines across diverse tasks, visual interfaces, and reasoning backbones. Operating post-SDPA and pre-W_O, SPLIT preserves stock fused attention, whereas attention-explicit baselines materialize attention tensors under eager execution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.