PathBalance: Diagnosing Path-Load Imbalance in Dense Language Models
Abstract
Dense language models expose no explicit routing distribution, yet fine-tuning can concentrate their internal computation in a small subset of attention heads, MLP neurons, and residual-stream dimensions. We study this phenomenon as path-load imbalance, measuring activation strength (A), gradient pressure (G), and activation-gradient attribution (C). On Intermediate Algebra, after the best checkpoint, the top 2% of MLP neurons increase their share of attribution load from 15% to 72%, and the top 2% of residual dimensions from 15% to 36%. Higher residual-C exposure predicts poorer retention of previously correct answers, and suppressing train-hot MLP neurons reduces accuracy from 32.7% to 1.5%. Building on these findings, we introduce PathBalance, a load-aware dropout method for full fine-tuning that measures the current load before every update and masks the top-2% hot units while the remaining network is trained with cross-entropy. Across three math benchmarks and five Qwen3/Qwen3.5 models, PathBalance improves over SFT by up to 16 percentage points on GSM8K and 50 on MATH, without changing the inference architecture. Masking the same number of randomly chosen neurons yields no gain, and a matched-epoch analysis shows that PathBalance shifts attribution load from overloaded neurons to previously cold ones. The code is available at https://anonymous.4open.science/r/PathBalance/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.