acceptodds
Under review as a conference paper at ICLR 2027

Utility-Augmented Transformers: Feedback-Conditioned Retrieval for Sequential Decision Making

Abstract

Offline reinforcement learning has advanced sequential decision-making. Within this framework, conditional sequence models such as Decision Transformer achieve competitive performance by predicting actions from trajectory histories and target returns-to-go. These models typically incorporate actions and rewards as input features without directly conditioning attention projections on feedback, which may limit their effectiveness on feedback-informative tasks, where identical observations can require different actions depending on what previous actions achieved. This paper examines when input-level feedback conditioning is sufficient and when it is not. We introduce the Utility-Augmented Transformer (UAT), in which shifted action-reward feedback modulates queries, keys, and values, and an exponentially smoothed utility state supplies persistent conditioning, with zero initialization reducing UAT exactly to its vanilla Transformer backbone. The additive pathway provably contains the linear input-conditioning subclass, while the multiplicative pathway realizes a third-order query-key-feedback logit that input-level injection cannot represent, yielding a persistent approximation floor for the latter under bounded stylized designs. The utility state also admits a compressed belief-state reading under explicit sufficiency conditions. Extensive experiments show substantial gains of UAT on controlled tasks with informative, timely feedback and competitive performance across standard offline RL domains, under comparable parameter budgets and without increasing asymptotic attention complexity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.