acceptodds
Under review as a conference paper at ICLR 2027

CAFRL: Causality-Aware Flow Based Policy For Online Reinforcement Learning

Abstract

Flow-based policies enhance maximum-entropy reinforcement learning by representing complex and multimodal behaviors through learned transport dynamics. However, entropy evaluation of implicit flow policies is generally difficult, motivating path-space regularization as an alternative to explicit density-based constraints. Existing approaches typically adopt isotropic exploration geometries, assigning identical kinetic costs to different action dimensions despite their heterogeneous relevance to tasks.We introduce CAFRL, a causality-aware regularization framework that incorporates action-to-reward effects into the exploration geometry of flow policies. CAFRL derives action relevance from interaction data through causal analysis and maps the resulting effects into a reference process, inducing anisotropic path kinetic regularization. This enables task-relevant directions to deviate from the reference process with adaptive path costs while maintaining a consistent regularization scale, thereby inducing a causal anisotropic exploration geometry. Across eight HumanoidBench tasks with five random seeds for each method, CAFRL achieves the highest average return compared with FLAC, ACE, and SimbaV2 after one million environment steps. CAFRL improves over the strongest baseline on each task by a median margin of 80.3%. These results support the empirical performance of the CAFRL configuration evaluated.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.