STODA: Structured Tiling via Orthogonal Design Adapters
Abstract
Modern foundation models increasingly rely on functional decomposition, representing complex mappings as mixtures of simpler primitives. Sparse Mixture-of-Experts (MoE) architectures provide an efficient realization of this principle, but often suffer from expert collapse. Existing load-balancing heuristics encourage statistical uniformity, yet do not enforce functional diversity by construction. We propose **STODA**: **S**tructured **T**iling via **O**rthogonal **D**esign **A**dapters, a framework that partitions the global mapping space using algebraic structures derived from Orthogonal Designs. Each expert follows a lightweight view–core–unview architecture: a fixed Hurwitz–Radon (HR) view transform, a learnable core module, and an inverse mapping. The anti-commutativity and skew-symmetry of the HR family provide a tractable algebraic mechanism for expert diversification: at the generator level they yield orthogonal token views and non-degenerate mixing, while for linearized cores they identify collapse conditions and balanced cross-term structure under routed mixtures. To validate the framework under strict resource constraints, we study STODA in the Parameter-Efficient Fine-Tuning (PEFT) setting for diffusion models. Empirically, STODA matches or surpasses state-of-the-art methods while requiring dynamic computation and fewer active parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.