Spatiotemporal Response Modeling for Diffusion Transformer Editing
Abstract
Controlling diffusion-transformer generation requires knowing where an internal edit propagates and how simultaneous edits interact. We introduce ST-Lens, a signed, coordinate-resolved response interface conditioned on residual block and denoising regime. It preserves spatial and video-time structure, revealing regional effects and cross-time interference. ST-Lens reduces response error by 37.15% against a frozen-path readout on 768 held-out fixed-dictionary interventions. We develop response-coupled residual editing (RCRE) to turn this information into coordinated temporal control: a full response Gram matrix jointly allocates a bounded residual write, and native-suffix feedback verifies its effect. Image control attains all 792 prescribed targets; Wan edits execute reversals, pauses and curved return trajectories. Response diagnostics motivate generation-policy candidates, whose final outcomes are measured with complete sampler rollouts. On ImageNet, a standard 50,000-generated-image evaluation per DiT arm against the 50,000 validation images gives 13.73% lower FID and 56.66% lower KID for scheduled guidance; a separate paired 1,000-image run gives a 1.312 sampling-and-decoding speedup. On MSR-VTT, advected sharing reduces mean FVD by 29.35% in a 1,024-pair cohort across three seeds and raises the reported three-seed means in all five VBench dimensions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.