Computation-Aligned Parameter Sharing Accelerates Grokking Beyond Weight Decay
Abstract
In grokking, a model memorizes its training data long before it generalizes, and closing the gap can take an order of magnitude more optimization steps. Accelerating grokking would therefore make generalization more compute-efficient. Weight decay is known to drive the transition, but it supplies only a global preference for smaller parameter norms and does not determine which training examples update the same parameters. We introduce ComRoute, a training incentive for computation-aligned parameter sharing on top of weight decay. Given labels for which examples share a reusable rule, ComRoute encourages those examples to route through overlapping experts in a sparse MoE model, so that their learning signals accumulate in shared parameters instead of redundant copies. Theoretically, in a stylized fixed-route model, we characterize generalization time in the small-initialization limit. Computation-aligned routing generalizes earlier than size-matched random routing whenever it strictly improves the threshold's critical learning rate; we also give two sufficient conditions under which a single ungrouped block is slower. Empirically, we show that on a factorized modular-arithmetic task with held-out factor combinations, a 2.3M-parameter MoE trained with ComRoute reaches 90% held-out accuracy after around 23K updates, 1.4 fewer than with size-matched random routing and 6 fewer than with weight decay alone. At 107M parameters and in OLMoE-1B-7B, ComRoute shows a similar overall trend. In short, weight decay helps a reusable solution win eventually, and ComRoute helps it win sooner.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.