Interaction Steering for Compositional Robustness in Vision Models
Abstract
Modern vision models are trained through outcome supervision: the predicted class should match the label. However, this supervision does not specify how visual factors should contribute jointly, allowing models to rely on spurious context or overlook useful combinations of object parts. We introduce Interaction Steering, a framework for mechanism supervision in vision models. Interaction Steering applies finite-difference operators to counterfactual edits of a model's score, isolating interactions such as object-background, object-co-occurring-object, and part-part effects. These operators audit how named visual factors jointly affect the score and define auxiliary steering and guard losses that move their interactions toward specified targets. Counterfactual edits are used only for these auxiliary losses, while the primary supervised objective remains on natural images. On Waterbirds, Interaction Steering increases worst-group accuracy from to on ResNet-50 and from to on DINOv2 ViT-B/14. On UrbanCars, it improves all three official shortcut gaps, including the combined background-plus-co-object gap. On PACO, it shifts the mean audited part-part interaction from to , while natural Top-1 changes by percentage points on a held-out PACO-LVIS test manifest covering five object categories. Further experiments retain Waterbirds gains with masks derived from the classifier itself and show larger interaction increases for the selected PACO part pair than for other audited part pairs. Interaction Steering enables direct supervision of how visual factors jointly affect a model's score, supporting both shortcut suppression and useful synergy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.