acceptodds
Under review as a conference paper at ICLR 2027

Attention Score without Multiplication

Abstract

Attention needs a query–key interaction, but that interaction need not be a dot product. We introduce Direct Attention, , which replaces coordinatewise products with subtraction, rectification, and accumulation while retaining softmax and value aggregation. We derive its exact row-softmax equivalence to coordinatewise minimum, its -distance-plus-key-bias form, and the conditions under which simpler separable scores lose query dependence. A budget-bounded ImageNet-1K conversion of a pretrained ViT-Small updates only Q/K weights and keeps every other parameter fixed. The resulting Direct checkpoint reaches 80.782% top-1 accuracy, compared with 81.392% for the frozen dot-product reference, a difference of −0.610 percentage points under identical official preprocessing. This is one adaptive development trajectory with repeated validation use. In a separate matched Sky130 HD synthesis study, the 16-bit Direct score unit occupies 0.115 times the dot-unit area, giving a Dot/Direct ratio of 8.67. The width sweep links this saving to the different arithmetic primitives; the SRAM-inclusive modeled total area falls by 0.42%. Together, the results identify query–key arithmetic as a practical co-design choice: a rectified difference supports accurate checkpoint conversion and a substantially smaller score unit, with system-level value determined by the surrounding memory and computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.