acceptodds
Under review as a conference paper at ICLR 2027

PRIM: Past-Action Regularization for Imitation with Flow Matching

Abstract

Behavior cloning commonly uses Flow Matching to learn distributions over future actions from expert demonstrations. Expert behavior can depend on temporal information absent from the current observation, making past actions a natural source of context for improving action prediction. However, strong temporal correlations in expert actions encourage policies to become highly sensitive to action history, allowing model-generated errors to be amplified when executed actions enter subsequent policy inputs. We introduce Past-Action Regularization for Imitation with Flow Matching (PRIM), which penalizes the velocity field’s Jacobian with respect to action history to reduce the sensitivity of generated actions to history errors. We show that strong conditional predictability from past actions requires history sensitivity, and derive the recursive propagation of action errors through environment dynamics and history inputs. Under sufficient conditions, controlling this sensitivity keeps accumulated errors bounded. Experiments on LIBERO-90 show that directly conditioning on past actions can lower success rates, while Jacobian regularization suppresses error accumulation and improves success.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.