Beyond Routing: Matrix Zonotopic Attention for Set Transformers
Abstract
Standard attention adapts token routing to the input while sharing its value projections across contexts. We study a complementary design axis for set learning: making feature transformations explicitly context-dependent and carrying transformation-derived features across layers. We introduce Matrix Zonotopic Attention (MZAttn), which parameterizes a transformation family using a learned center matrix and generator matrices modulated by bounded, permutation-invariant context gates. Rather than using this family only to produce a single-vector update, MZAttn propagates separate center and generator feature streams and connects them to prediction through a generator-aware readout. To characterize operator variation, we introduce Transformation Degrees of Freedom (TDOF), which counts the independent matrix directions along which a specified operator varies with context. At the attention-core level, we establish exact representations for specified gated operator families and quantify approximation limits within fixed matrix-bank classes. We further show that identical signed and magnitude summaries can become distinguishable after the same rowwise normalization. For the normalized magnitude readout, we characterize local information capacity and construct a complementary two-row code supporting stable local recovery and scalar linear prediction through a finite head, with an explicit approximation bound. The resulting architecture preserves the permutation structure required for set prediction. We investigate its inductive bias through set-learning experiments. Together, these contributions frame set attention around three complementary choices: where to route information, how to transform it, and what transformation information to retain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.