Early-Exit-Aware Partitioning for Distributed DNN Inference
Abstract
An early-exit strategy allows samples to exit when a certain confidence threshold is reached at intermediate layers of the model. However, distributing early-exit networks across multiple machines introduces unique complications. Applying naive layer-wise partitioning can negate the advantages of early exits: unlike conventional model parallelism, where all inputs traverse every partition, the volume of data reaching downstream machines varies. This produces an uneven workload across the cluster. Addressing this imbalance requires jointly deciding where to place early exits and how to partition and distribute the network across nodes. In this paper, we propose a load-balanced, distributed, model-parallel inference strategy that allocates compute capacity in proportion to the inference load at each stage. We evaluate our approach across multiple CNN architectures (DenseNet-121, ResNet-50, and AlexNet) on standard benchmark datasets (ImageNet, CIFAR-100, and CIFAR-10). Experimental results demonstrate that our architecture significantly reduces inference latency and outperforms conventional model partitioning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.