Harmful Partial Interventions and Benign Leakage in Concept-based Models
Abstract
Artificial intelligence systems deployed in high-risk settings increasingly require models that support meaningful human oversight. Concept-based models (CMs) offer a promising framework by enabling users to inspect and intervene on intermediate, human-understandable concepts. However, concept interventions can counterintuitively worsen downstream predictions, a failure mode we formalise as intervention harm. We distinguish harm after all concepts have been intervened on from harm under partial interventions, the more likely practical use case. Modern CMs often achieve high accuracy by encoding, or leaking, additional task-relevant information in their intermediate concept representations. Recent work argues that such leakage can be benign with respect to intervention harm when concept representations satisfy two conditions: sufficiency and localisation. We prove that these criteria are not sufficient to guarantee that concept interventions improve downstream performance. Across controlled and real-world settings, we find that concept supervision incompleteness associates with increased intervention harm. Importantly, optimising the model when all concepts are intervened on, a previously proposed leakage mitigation objective, reduces endpoint harm but does not reliably prevent harmful partial corrections. Hence, we introduce a complementary monotonic intervention objective that directly penalises increases in task loss along partial intervention trajectories. Combining endpoint and monotonic intervention training substantially reduces the prevalence and magnitude of partial intervention harm while preserving concept and task predictive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.