Gradient-Cancelled Concept-Based Models
Abstract
Concept-based models are designed to support human intervention: an expert can correct an incorrect concept at test time and improve task prediction, a capability particularly important in high-stakes domains. Joint training typically achieves the strongest task performance, but this can come at the cost of information leakage, whereby concept representations encode task-relevant information not captured by their intended semantic meaning. We provide an information-theoretic characterisation of task-driven leakage in joint training: under task-head convergence, the encoder update contains an explicit mutual-information maximisation term that encourages concept representations to capture task-relevant information beyond the annotated concepts. We introduce Gradient-Cancelled Concept-Based Models, which use a Gradient Reversal Layer to cancel the task-driven encoder gradient, leaving the concept encoder to learn from concept supervision alone while training the task head concurrently. Our model requires a simple training-time gradient modification and applies to any concept-based model where task gradients reach a shared encoder via concept representations. Our experiments show that Gradient Cancelled (GC) training results in significant task improvements upon intervention over vanilla concept-based models, while maintaining strong concept and task baseline performance, making these models better suited to in the human-in-the-loop settings they were originally designed for.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.