The Credal Information Bottleneck
Abstract
The information bottleneck (IB) finds a compressed representation that preserves information relevant to a task. Traditional IB defines this compression–prediction tradeoff for a fixed joint input–output distribution. While robust IB variants account for perturbations in that distribution, they do not explicitly preserve the information needed to estimate the uncertainty at the output of the bottleneck. To address this fundamental limitation, we propose the credal information bottleneck (CIB), which for the first time represents uncertainty in the data-generating distribution and the data processing, and propagates it over the bottleneck, enabling its estimation and categorization at the bottleneck output. Importantly, CIB disentangles aleatoric, data-induced and processing-induced uncertainty. This knowledge about the sources of uncertainty allows performing specific and selective interventions at deployment-time. To do so, CIB leverages the formalism of credal sets, i.e., convex combinations of probability distributions, to characterize and propagate the uncertainty and the information about the uncertainty sources through the bottleneck. By replacing point-valued representations with credal sets, CIB fundamentally changes the very definition of information object and information distortion. We introduce variational CIB (VCIB) to make CIB differentiable, and prove that the VCIB objective upper-bounds the CIB up to an additive term. We further prove that every pooled-efficient IB representation incurs a strictly positive credal-distortion floor, which CIB provably overcomes. On three colored benchmarks with environment conflict, VCIB improves post-deployment worst-case accuracy by up to 6.4 percentage points over the strongest baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.