Robust Gated Linear RNNs via a Single Linear Time-Invariant Recurrent Layer
Abstract
Gated linear Recurrent Neural Networks (RNNs) incorporating input-dependent gate mechanisms have shown promising performance on sequence modeling. While input-dependent gating allows their dynamics to adapt flexibly to each input, it also makes the dynamics sensitive to task-irrelevant inputs. To characterize and address this issue, we directly analyze the gate mechanism and show that the gate has a *non-constant rate of change*: When the gate value is close to one, a decrease in the pre-activation causes a larger change in the gate than an increase of the same magnitude. This asymmetry can cause a downward shift in the gate values, making it difficult to preserve information over long sequences. Moreover, a larger learning rate can amplify this downward shift since the update of the pre-activation scales with the learning rate. This insight leads to our proposal: We replace a single input-dependent recurrent layer with a *linear time-invariant* (LTI) *recurrent layer* in a gated linear RNN while leaving the rest of the architecture unchanged, which we call the *single-LTI hybrid*. Since the state-transition matrix of the LTI layer is independent of the input, this replacement mitigates the effect of task-irrelevant inputs and large learning rates. Interestingly, replacing only a single layer is sufficient because residual connections commonly used in gated linear RNNs provide an identity path, allowing the features extracted by the LTI layer to be preserved through the subsequent layers. Our experiments show that the proposed single-LTI hybrid improves trainability on long-sequence modeling tasks, particularly with large learning rates, reduces undesirable sensitivity to inputs, and achieves competitive or better performance on language modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.