acceptodds
Under review as a conference paper at ICLR 2027

PRECISION-WEIGHTED ATTENTION: RESTORING UNCERTAINTY TO THE TRANSFORMER FOR SEQUENTIAL RECOMMENDATION

Abstract

The Transformer treats every token with uniform confidence: attention weights a position by how relevant it is to the query, never by how much evidence supports its representation. Sequential recommendation violates this assumption in every attention row, where an item observed a handful of times sits beside one observed thousands of times. We read the pre-norm Transformer layer as one cycle of a Bayesian filter: attention observes, the residual connection updates, and the feedforward sublayer predicts. Under this reading exactly one quantity is missing from the standard layer, the precision of its state. Restoring it gives precision-weighted attention (PWA): each token carries a positive precision alongside its hidden state, and one observe–update–predict cycle per layer uses it as a bias on the attention logits toward better-supported positions, as the gain of the residual update, and as a quantity transported through the feed-forward sublayer by its Jacobian. The observation precision is closed-form, computed from the dispersion of the pooled values and their effective sample size, and PWA adds 132 parameters, four of which train, with no learned gate. On SASRec, BERT4Rec, and HSTU across six datasets, under one recipe fixed for every backbone and dataset, with 20 matched seeds per cell (720 runs), PWA improves mean NDCG@10 in 17 of 18 backbone– dataset pairs, 15 surviving Holm correction, with the largest gains on rare items and short histories. A matched-capacity control and component ablations locate the gain in the observe and update steps, and SASRec+PWA outperforms STOSA, the closest uncertainty-aware recommender, on every dataset. The filter reading is a structured approximation, not a calibrated posterior; we state each simplification and every negative cell. Precision as a first-class state modifies the Transformer itself, and its extension beyond recommendation is future work

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.