acceptodds
Under review as a conference paper at ICLR 2027

Unnormalized Relaxations for Robust Modular Learning

Abstract

Modular learning scales Large Language Models by routing inputs through frozen, domain-specific experts rather than retraining a monolithic model. The robust gating framework of Cortes et al. (2026) achieves minimax optimal performance under arbitrary shifts in the mixture proportions of the known source domains, but its strict global normalization (: ) couples the gate to every expert's predictions, creating a costly synchronization bottleneck at training time. We prove that this constraint is not only computationally unnecessary but *yields less favorable provable guarantees*. We introduce two *unnormalized* gating families and develop a complete theoretical, algorithmic, and empirical framework around them. The first bounded family replaces the data-coupled mass constraint with an expert-independent structural ceiling (), depending only on the gating function . We prove that achieves a provably tighter minimax robust guarantee in specialized regimes via a novel *Normalization Bonus* , with a matching lower bound; an ordering of empirical Rademacher complexities under mild conditions ( for ); *Exact Rejection Sampling* for unbiased inference with a structural efficiency bound; and a *Distillation Inheritance* theorem proving that causal student routers preserve the teacher's minimax bounds up to a controlled gap, with no causal projection penalty and an assumption-free total variation variant. The second locally-normalized family enables fully pointwise updates, with robust minimax guarantees via a *Selection Gain* and parallel distillation guarantees. Experiments on 1B+ parameter LLMs and 20-expert code generation ensembles confirm the theory: unnormalized gates reduce wall-clock time by 50% while strictly improving NLL over both monolithic and baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.