Causal Analysis of Physical Understanding in Visual Representations
Abstract
Decoding physical variables from visual representations does not establish how their encoded information contributes to task prediction. The challenge is to separate target-specific contributions from general intervention effects and the loss of information shared with related variables. We introduce CausalProbe, a framework for representation-level causal analysis. It progressively erases target-predictive directions while protecting selected readouts of related variables. Random and reverse-role erasure provide comparisons for interpreting the resulting task effects. Across six frozen encoders, erasing second-order geometry impairs grasping-affordance prediction on SuctionNet much more than erasing first-order geometry, supporting its stronger causal contribution. On Physion++, spatial erasure similarly affects future-contact prediction more than motion erasure. Removing either family without protection also reduces the other’s predictability, indicating overlap between spatial and motion representations. Our framework moves beyond identifying what visual representations encode to examining how that information supports task prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.