Low Correlation Is Not Separation: What Gradient Isolation Delivers for Epistemic and Aleatoric Uncertainty
Abstract
Uncertainty estimates should distinguish prediction error that can be reduced through learning from ambiguity in the available labels. Predictive decompositions do not generally identify these two sources, and their estimates can be strongly correlated. We study structural separation in supervised latent variable models, using distinct parameter heads and supervision signals for error-related and ambiguity-related uncertainty, instantiated as a Credal Concept Bottleneck Model. We state sufficient conditions for isolating the uncertainty-head gradients, verify them with a loss-by-parameter audit of a trained model, and distinguish this property from output decorrelation and semantic validity. We evaluate these properties on five benchmarks spanning annotator and corpus supervision. Low EU–AU correlation is common, but it is informative only on CEBaB, where both heads track their targets: without a decorrelation penalty the scores are weakly correlated ( and for two encoders), and the error-related score matches or exceeds the mutual information of a three-member ensemble (AUROC ) from a single model. On HateXplain and GoEmotions the ambiguity head collapses to a constant, and on two QA benchmarks the error-related score is at chance. On every benchmark, the model's own maximum class probability detects errors better. Removing ambiguity supervision collapses the ambiguity head, and an explicit decorrelation penalty collapses one head in every seed. Gradient isolation and low correlation are therefore achievable, but neither establishes a useful decomposition; we identify the target properties that decide the outcome and propose what an evaluation of separately supervised uncertainty must
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.