acceptodds
Under review as a conference paper at ICLR 2027

On Reward Design for Aligned Decision-Making under Uncertainty

Abstract

In human-AI cooperation, a utility function elicited from a human decision maker is routinely repurposed as a reward for the artificial agent. When the agent communicates set-valued predictions, extending the original utility to a reward over imprecise information is nontrivial, as the expanded utility has many more degrees of freedom than the original. Existing methods assume that the value of an imprecise piece of information reduces to an aggregate over its constituent precise states, an assumption that is frequently violated. We propose a general framework that makes no such reduction. Decider's behavior under uncertainty is explicitly modeled via selection strategies induced by decision criteria. We introduce the notion of validity, which ensures that a criterion respects the rationality encoded in the utility function, and the stronger notion of alignment, which guarantees that the resulting reward faithfully preserves the original preferences. We prove constructive compatibility results between classical criteria (Laplace, maximin, maximax, minimax regret) and standard aggregators (min, max, avg), and provide an algorithm to build aligned rewards. Our framework accommodates arbitrary sets of terminal acts, bridges the gap between utility elicitation and reward design for set-valued information, and generalizes naturally to richer uncertainty representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.