CausalCache: Aligning Computation with Action in Diffusion Policies
Abstract
Accelerating diffusion robot policies requires identifying computation that can be reused without compromising control. Feature similarity offers a convenient signal, but overlooks how approximation errors are transformed before they reach the robot: the remaining denoising network can absorb or amplify a local perturbation, while robot kinematics can turn similar joint-command errors into different end-effector displacements and rotations. We therefore view cache allocation as preserving task-space fidelity rather than matching intermediate features. We introduce CausalCache, a training-free framework that compiles cache policies from these propagated task-space consequences. Single-block interventions recompute the full network suffix, and an end-effector pose (EE-Pose) objective measures the resulting position and orientation errors. A fixed-budget compiler minimizes the tail risk of intervention costs using , while conditional pair refinement evaluates interacting changes in the context of the compiled policy. Calibration requires neither policy retraining nor task-success labels; deployment executes a frozen step-layer matrix with no online selection overhead. On RoboTwin2, CausalCache retains near-full task success with block reuse, delivering action-denoising and complete-policy inference speedups. Eight calibration episodes perform comparably to forty, while SmolVLA experiments on a physical SO-101 demonstrate the advantage of task-space scoring across three manipulation tasks. These results establish propagated action consequences as a practical basis for allocating computation in diffusion robot policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.