IAR: Implicit Action Repetition for Policy Optimization with Sequence Critics
Abstract
In actor–critic learning, policy updates depend on both the learned value function and the action contexts in which it is queried. We study action repetition as a mechanism for constructing these contexts during policy optimization. We introduce implicit-action repetition (IAR), an actor-update mechanism for continuous control with sequence-conditioned critics. For each replay state, IAR samples an initial action and constructs prefixes that repeat it for different durations. The policy then samples a continuation action conditioned on each prefix, and the actor is optimized using the resulting sequence values. Repeated prefixes and critic parameters are held fixed during differentiation, allowing the actor to learn under hypothetical sustained-action contexts without requiring those repetitions to be executed in the environment. Once the initial action is sampled, all prefixes are available, allowing a causal Transformer to evaluate the continuation distributions in parallel. This removes sequential feedback among continuation predictions during actor updates, while environment interaction and critic bootstrapping retain autoregressive sampling. Across all 50 manipulation tasks under the Meta-World ML1 protocol, the resulting method achieves an aggregate success rate of approximately , compared with approximately for T-SAC. These results support repeated-action conditioning as an effective mechanism for policy optimization with sequence critics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.