Beyond Stability: Attention-to-MLP Interaction in Infinite-Depth Transformers
Abstract
Standard Transformer blocks pass the attention output to the MLP. However, we find that the depth scaling used to stabilize residual accumulation can make this interaction vanish as depth increases, even though the architecture remains computationally sequential. We study this effect under Depth-P and CompleteP residual scalings at initialization and during training (one gradient-descent update). By jointly controlling representation moments and weight-reuse responses, we identify sufficient stability conditions under which the interaction coefficient can remain order one independently of residual scaling. As a result, the accumulated interaction can remain order one as depth increases. By contrast, under the standard tied choice, the accumulated interaction decays as for Depth-P and for CompleteP, where is the model depth. Consistent with this analysis, experiments support the predicted depth trends and show that suitable independent coefficients can retain interaction and improve learning over the tied baseline. However, we also observe that the performance gains are sensitive to the interaction coefficient. This leads us to explore input-dependent gates that adapt the interaction strength during training. In our comparisons, these adaptive gates indeed further improve performance, with signed SiLU and linear gates outperforming positive-only sigmoid gates. These findings motivate separating the interaction coefficient from residual scaling, enabling a model design choice between retaining interaction when beneficial for learning and using explicitly parallel blocks for concurrent computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.