Factored Diffusion Policies: Compositionally Generalized Robot Control with a Single Score Network
Abstract
Robotic tasks are typically specified by a tuple of factors, such as the goal to be reached, the obstacles to be avoided, and the appearance of the scene, and collecting expert demonstrations for every combination of factor values grows combinatorially. We present factored diffusion policies: a single diffusion network trained with per-factor null-token dropout, whose score decomposes additively across factors. When the factors are approximately conditionally independent given the action and observation, the composed score approximates the joint score uniformly over all factor combinations, so a training set in which every factor value appears at least once suffices: the task budget grows with the sum of the factor cardinalities rather than their product. The decomposition residual equals the action-gradient of the factors' interaction log-ratio and is measurable from forward passes, so the assumption can be checked per task, and the same structure yields a closed-loop trajectory-tube certificate. We evaluate our approach using drone racing simulations where a drone must navigate through gates while balancing speed and path curvature against tracking overshoot which can cause gate crashes. Trained on such a cover, the factored network passes several times more held-out gates than the same architecture without dropout, replicated across disjoint covers, seeds, and two and three factors, while the same composition rule over separately trained networks generalizes far worse. Injected factor interaction raises the measured residual and degrades composed inference, and our experiments indicate that the gain comes from the factored training, and the decomposition supplies the task budget, the certificate, and the approximate conditional independence check.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.