acceptodds
Under review as a conference paper at ICLR 2027

WAVE: World-Model-Augmented Value Learning for Action-Chunk Reinforcement Learning

Abstract

Action chunking provides a temporal abstraction for long-horizon reinforcement learning, but a chunk critic must infer what an entire H-step action sequence will accomplish from reward and temporal-difference supervision alone. A world model can predict these consequences, but substituting its predicted endpoints into the Bellman backup lets model error change the preferred chunk: we construct an example in which a small endpoint error yields a suboptimal policy despite exact rewards and an exactly solved Bellman equation. With observed endpoints, by contrast, the Bellman fixed point is optimal for any world model, so model error can affect only what the critic represents. We therefore introduce WAVE (World-Model-Augmented Value Learning), which conditions the critic on a causal world-model posterior through a zero-initialized residual adapter and trains its features to predict the model's current latent state and the H-step consequence of the same chunk, while every Bellman target uses real rewards and observed endpoints. On two OGBench manipulation tasks with three seeds each, WAVE improves on Q-Chunking, raising the area under the success curve over 1M online interactions from 0.832 to 0.929 on Cube and from 0.922 to 0.956 on Scene and reaching 70% and 80% success with 76–85% fewer real interactions on average; substituting predicted endpoints from the same world model instead yields near-zero success. Ablations attribute most of the auxiliary objective's gain to predicting future rather than current latent states.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.