acceptodds
Under review as a conference paper at ICLR 2027

AutoValueWorld: Learning Value Dynamics for Consequence-Aware Decision Making

Abstract

In multi-turn interactions, language model assistants can shape users' deliberation and the conditions for subsequent decisions. Responses with similar immediate ratings may support or violate human values differently, changing the relative utility of later actions. Response-level judgments alone don't specify how these value-relevant states evolve under successive responses. We introduce AutoValueWorld, an action-conditioned value world model that separates stable value importance from a dynamic, ten-dimensional value-support state. Training combines procedural supervision of state transitions and simulated reactions with human-rating supervision from PRISM. Candidate-difference supervision learns relative value consequences and recursive state prediction supports a frozen gain-based selector between one-step and three-step policies. Experiments show candidate-difference supervision reduces held-out candidate-contrast RMSE by 13.1% and 13.8% at 3B and 7B. In a study of 5,000 procedural episodes, selective lookahead improves mean realized return over one-step selection for value alignment at 3B and weighted reaction at 7B. These results support explicit value-dynamics modeling and selective lookahead for consequence-aware decision making in procedural environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.