Conditional Complementary Feature Learning for Early-Exit Networks
Abstract
Early-exit networks enable adaptive inference on resource-constrained devices by allowing easy inputs to terminate early while harder inputs receive additional computation. Conventional depth-wise designs attach classifiers to a shared backbone, introducing two limitations: architectural constraints, which tie early-exit capacity to backbone position, and early-late interference, where shared features must support both immediate classification and continued processing, potentially degrading early-exit accuracy under joint training. Our approach is motivated by the intuition that shared-prefix computation may contribute to these limitations. Our key insight is to separate architectural coupling from sequential execution. Rather than evaluating a single network to increasing depth, we organize inference as a sequence of complete, exit-specific subnetworks. Since earlier representations are already available when a later exit is reached, later computation can focus on learning features that complement them. Guided by this principle of conditional complementary feature learning, we introduce an early-exit framework in which each subnetwork processes the original image through a full-depth feature hierarchy. This permits narrow, full-depth early predictors and relaxes the requirement that earlier features preserve all information needed for later computation. Each later subnetwork conditions feature extraction on the stage-end features of all preceding subnetworks, while its classifier combines the current and earlier final embeddings. Earlier representations thus guide subsequent extraction while remaining directly available for prediction. We instantiate the framework with CNN backbones and evaluate it on ImageNet and embedded CPU and GPU platforms. Our models improve the accuracy-efficiency trade-off, requiring up to 2.2× fewer FLOPs than the compared early-exit baselines at matched accuracy. Comparisons with Dynamic Perceiver show gains of up to 10.2 and 14.0 percentage points in top-1 accuracy at matched FLOPs and latency, respectively, and up to 2.5× lower latency at matched accuracy. Our work offers a new perspective on early-exit design, in which additional computation is conditioned on, and complementary to, what has already been computed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.