Modulated Transformer
Abstract
Softmax based attention suffer from several challenges like attention sink and attention noise. A well known reason for which is that all attention scores need to sum to one. Hence, a natural remedy would be to explore alternatives to softmax. A recent advancement that replaces softmax is the DIFF attention which relies on correcting a base information carrying softmax attention map by subtracting another sequence dependent softmax attention map instead of per token-position dependent gate as in gated attention. We introduce MODULATED TRANSFORMER that computes a sequence dependent correcting mechanism by modulating the amplitude of a base softmax attention map signal by taking hadamard product with the second correcting attention map. Modulation of a base softmax signal by a suitable modulating function results in sharper and sparser final attention patterns. At 780m scale, Modulated Transformer surpasses softmax based Transformer in language modeling and achieves significant model size and token efficiency as indicated by a lower perplexity score with only 58% of the model size of 1.3B Vanilla Transformer with similar token budget. It also outperforms softmax based Transformer by 3.15 average accuracy points and by 3.24 points after pruning attention weights and is competitive with next strongest baseline DIFF on short-context language modeling tasks. With these results, the Modulated Transformer advances the abilities of existing language models and opens up a new avenue for research on developing new expressive and stable modulating functions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.