Depth Separation for Learning Hierarchical Interactions with Nonlinear State-Space Models
Abstract
As sequence models emerge as efficient architectures for long-context modeling, it becomes important to understand the role of depth in recurrent architectures which combine recurrent memory with nonlinearity across layers. In this paper, we study the effect of depth through balanced-tree parity targets and show depth seperation phenomenon on Boolean input sequences. First, under norm constraints and bounded activation regularity, we obtain the lower bound of approximation error for deep nonlinear SSMs which means every constant-depth model cannot represent this target regardless of its hidden dimension or nonlinear width. Second, we obtain an upper bound for a logarithmic-depth SSM learned via a normalized layerwise coordinate-descent procedure when fitting this target from polynomially many noiseless samples, under the norm constraints assumptions and with constants independent of the sequence length. Third, we construct a depth-two nonlinear SSM that exactly represents a deep quadratic target, which cannot be represented by any constant-depth fully connected network of comparable size, demonstrating recurrent memory can substantially reduce the nonlinear depth needed for exact representation. These results characterize depth separation through the interplay among memory dimension, nonlinear width and network depth while extending the theoretical understanding of deep nonlinear recurrent architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.