One Size Doesn't Fit All: Depth-aware Manifold-Gated Mixture-of-Experts for Deep Imbalanced Regression
Abstract
Existing representation-based Deep Imbalanced Regression (DIR) methods typically enforce label–feature geometry at a fixed representation depth. This one-size-fits-all strategy can overlook depth-dependent variation in label–feature alignment. We propose Depth-aware Manifold-Gated Mixture-of-Experts (DMG-MoE) to address this depth-dependent variation. It measures Depth-wise Manifold Alignment (DMA) across selected backbone depths by comparing an input's soft feature neighborhood to its soft label neighborhood via KL divergence, and uses this signal to guide the allocation of prediction weight across depth-wise experts. We distill geometry-aware supervision from a training-time DMA teacher into a label-free student gate that produces routing weights without label access at inference. We parameterize routing via a stick-breaking scheme that sequentially allocates prediction mass across depth-wise experts. This enables sample-specific, geometry-guided depth adaptation. On AgeDB-DIR and IMDB-WIKI-DIR, DMG-MoE achieves the lowest overall MAE and GM among the compared methods. We also observe gains on the tabular benchmarks (Abalone, Airfoil) and with two additional backbones (ResNet-152, ConvNeXt-Tiny). The learned routing aligns with its DMA teacher signal, supporting the geometry-aware mechanism.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.