MoENet: Hidden Representations as Layer-Wise Experts for Adaptive Aggregation
Abstract
Network depth remains a pivotal yet sensitive design variable in deep learning, which intermediate layers can yield representations superior to the final layer, while excessive depth invites overfitting and optimization pathologies. We present MoENet, which systematically harnesses hidden representations across depth by treating each layer as a distinct expert and adaptively aggregating them via input-aware routing weights. We further introduce a sample-dependent softmax temperature, so that the sharpness of expert selection is modulated on a per-input basis. Extensive experiments demonstrate consistent gains across diverse backbones (e.g., ResNet, ConvNeXt) at only negligible additional computational and parameter overhead. Beyond accuracy, MoENet markedly exhibits stronger resilience to overfitting at extreme depths. Finally, incorporating MoENet into the representative baselines of seven additional image recognition benchmarks consistently yields significant improvements, underscoring the generality and broad applicability of our approach.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.