Learning Individual Representations from Aggregate Supervision
Abstract
Aggregate-supervised learning observes a label for a bag of instances while the instance labels remain hidden. When that label is a sum of finite-support contributions, an instance predictor can be trained by maximizing the probability of the observed total. We present FS-Conv, a likelihood layer that separates the instance predictor, the map from predictions to contribution probabilities, and the convolution backend. Under conditional factorization, discrete convolution computes the exact model-implied aggregate probability mass function. With Bernoulli contributions, the objective reduces to established count likelihoods; categorical contributions support integer-valued sums, and known signs encode cancellation. We compare this objective with mean matching and a Gaussian approximate likelihood. In three-seed digit-sum experiments on MNIST and on SVHN with an ImageNet-pretrained ResNet-18, FS-Conv has lower mean aggregate error and discrete aggregate negative log-likelihood, and higher hidden-digit accuracy, than both alternatives. Random-sign MNIST also favors FS-Conv on the reported means. On feature-defined Criteo click-count bags, FS-Conv is competitive with the evaluated label-proportion baselines. These results favor exact discrete likelihood training on the digit-sum and known-sign tasks; instance recovery still depends on the representation and the adequacy of the factorized model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.