An Attention Free Alternative to Language Modeling
Abstract
Most of the recent Transformer models have attention mechanism as a key component. Attention mechanism is the de-facto standard of transformer model as it enhances performance significantly. Nevertheless, it is well-established that attention mechanisms entail significant computational expenses. This places a restriction on model and input size when computing is limited. Additional limitations of attention mechanisms include the attention sink phenomenon. The primary goal of language model is to have efficient sentence representation. In attention mechanism this goal is achieved through similarity computation of tokens based on a starting representation, which is subsequently refined across different levels. We employ an algebraic approach. Tokens are mapped to a lower dimension first. Token pairs are subsequently generated based on the specified attention window size. The two-dimensional space spanned by the token pairs are encoded through a similarity based encoding approach. Our proposed method, the Algebraic language model (ALM), was evaluated using Wikitext2 data. Our experimental results using a laptop demonstrate improvement of the perplexity score by more than 25% in ALM relative to conventional Transformer models. Beside, ALM has substantially less computational complexity compared to Transformer models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.