Variational EM Energy Models for Weak Multi-Instance Multi-Label Learning
Abstract
Multi-Instance Multi-Label (MIML) learning trains models on bags of instances, where each bag may associate with multiple labels. However, the presence of noisy instances (defined as weak instances) in bags probably leads to mislabeling and missing labels (referred to as weak labels). We summarize the framework of excluding these weak situations and identifying ground-truth labels in MIML as weak multi-instance multi-label (wMIML) learning. Existing methods address either label noise or instance noise in isolation and do not explore the interplay between noisy instances and noisy labels. Methods that jointly handle data noise and label noise are mostly developed for single-instance data such as images and can hardly cover the multi-instance scenario. In this paper, we address wMIML from an energy-based perspective, casting the removal of both noise sources as a principled inference problem. We propose a variational-EM energy model in which each instance carries a latent support assignment over concepts. The energy measures the coherence between instances, their assigned concepts, and the observed bag labels, so that the two noise sources manifest as violations of one shared notion of coherence and are denoised through a single assignment. Variational EM on this model derives an approximate evidence lower bound from first principles, and the constraints induced by this bound couple label denoising and instance selection through the shared support assignment, so that the two tasks inform each other within a single bound. On ten MIML benchmarks under joint instance and label noise, VEM achieves the best F1 in 20 of 30 settings and stays within the top two on 26 of the 30, its only exceptions being the single-concept benchmark (FashionMNIST at flips 0.2 and 0.3) and two ceiling-saturated Scene settings. It consistently exceeds every baseline in the joint-noise regime where the two degradations must be handled together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.