acceptodds
Under review as a conference paper at ICLR 2027

WHEN DOES EXACT-ZERO ATTENTION HELP?

Abstract

When does exact-zero attention improve prediction over bounded dense softmax? We answer this by disentangling three mechanisms: bounded representation, implicit optimization bias, and network-level nuisance propagation. Theoretically, balancing a lower bound on bounded softmax residual-noise risk against a finite-sample support-recovery bound for sparsemax yields a sufficient condition: exact zeros reduce excess risk whenever support-estimation error drops below the retained noise floor. Across 2,400 controlled fits, this trade-off spans an empirical phase map over support sparsity, noise scale, and sample size. Solving both model classes to exact convex optimality decouples representation from optimization. The sparse advantage and its contraction with score range persist without an optimizer, while a closed-form leak term predicts the gap between optima across twelve orders of magnitude. Inside 300 frozen forecasters, the network-level mechanism is isolated: perturbing only the positions zeroed by sparsemax shifts predictions to of what an equally large random set does, against for its softmax twin. On real data, an input coordinate estimated strictly from the training split tracks the regime across the settings we test. Over 798 configurations drawn from a 1.2-million-record cryptocurrency panel and a 121-series US macroeconomic panel, spanning both attention directions, exact zeros lower test error in 773 cases (): by across the cryptocurrency cross-section, along a size-matched time axis, and across the macroeconomic series, with intervals robust to overlapping-window dependence. Finally, a decision rule frozen and registered in advance returned all four of its predicted wins on two untouched datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.