EMBODIED XAI: TRACING HIERARCHICAL AGENT DECISIONS VIA 4D SPATIO-TEMPORAL SALIENCY
Abstract
Robots steered by vision-language models are being proposed for spaces people occupy, where an unaccountable choice carries real cost. Yet when these physical AI agents silently pick a target, per-frame 2D saliency cannot say what in the physical world drove them: the planner is re-queried continuously while the robot drives, so each map sits in its own moving camera frame, leaving thousands of disconnected images that cannot be pooled, sliced, or compared. We present 4D spatiotemporal saliency, a model-agnostic framework that relocates attribution from the camera to the scene: an attribution stage computes importance scores with respect to the planner's sampled steering token, each attributed pixel is lifted into world coordinates through depth and pose, and the results accumulate across queries and trials into one heat field growing through time. The explanation thereby becomes an object that can be added up, grouped by outcome, differenced against a reference, replayed, and physically intervened upon. We instantiate it with chosen-action integrated gradients on a quadruped driven by ambiguous instructions, and use the field to certify scene-level attribution patterns as repeatable, separate the signatures of success and failure, watch a decision harden, and run closed-loop occlusion experiments that persistently target the same physical region as the camera moves.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.