acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Discovery of Motion-Effect Circuits in Video Diffusion Transformers

Abstract

Video diffusion transformers make camera and object motion observable, but observation and attribution do not by themselves identify which internal computations support a particular motion effect. Large attribution can reflect denoising shared across motions rather than support specific to the requested direction. We introduce CAP-MED, a counterfactual activation-patching protocol for motion-effect node discovery. It uses target-vs-other attribution to select compact node sets across denoising steps, layers, and submodules, then tests whether exchanging their activations recovers and suppresses the effect between matched motion-instructed and empty-action rollouts. On Wan2.2, specificity-selected sets achieve median normalized recovery of 0.99 and suppression of 0.97, compared with 0.15 and 0.59 for attribution-magnitude selection at the same node budget. Selector controls and independent camera and object readouts support the motion-specific interpretation of these effects. Further head-level interventions identify conditional dependence between selected sets and a value-dominant effect within selected camera-pan heads. LTX-Video experiments test the full workflow on a second architecture and extend it to color temperature. The evidence characterizes selected node sets and their functional roles rather than a complete motion-generation circuit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.