Phaseformer: Multimodal Policies without Iterative Compute or Stochastic Inference
Abstract
Deterministic behavior cloning policies struggle to capture multimodal robotic action distributions, suffering from mode averaging. Accepted solutions transform stochastic samples into robot actions, often via iterative computation (e.g. diffusion and flow matching). Flow matching backbones themselves are deterministic however, leading us to ask: why do flow matching models succeed at capturing flows between endpoint distributions while deterministic policies fail in similar situations? In this work, we prove that flow matching is a special case of behavior cloning, show that conditioning a deterministic policy on task progress lets it reproduce multimodal demonstrations whose start states vary, and characterize a score-based regularizer that pulls the policy toward the locally dominant demonstrated velocity where modes overlap. We demonstrate this empirically on multimodal robot navigation and in-plane manipulation tasks. Our findings indicate that lightweight, single-step deterministic policies can solve multimodal tasks while providing a Pareto-optimal trade-off for ultra-low latency, CPU based control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.