acceptodds
Under review as a conference paper at ICLR 2027

Operator-Valued Q-Learning: Compositional Bellman Recursions for Order-Sensitive Decision Making

Abstract

In quantum pulse sequences, spacecraft attitude control and in-hand reorientation, actions act on a hidden physical state as non-commuting operators, so the outcome of a plan depends on the order of its actions. Standard value functions either observe this latent state and learn a nonlinear function of it, or record the action history, which in tabular form costs entries for actions and horizon . We instead let the value function carry order: each is an operator, and the Bellman backup applies the action's operator to the continuation value, . Unrolled along a plan, the recursion builds the ordered product of the actions' operators, and the value at any latent state is a linear trace readout. For policies that do not observe the latent, this representation is exact; the backup contracts for unitary and orthogonal dynamics, and an asynchronous operator temporal-difference (TD) scheme learns it with almost-sure convergence. Control needs a maximum over actions, which operators lack: by Kadison's anti-lattice theorem, self-adjoint operators generally have no least upper bound. An operator log-sum-exp instead gives one operator per state–action pair whose readout upper-bounds the optimal value at every latent state, reducing to soft -learning when actions commute. We also prove that scalar values on observations that ignore action order must confuse swapped histories, that exact history tables grow exponentially for free action groups but polynomially for nilpotent ones, and a Hankel-rank lower bound on the operator dimension. On random and -pointing control tasks, operator -learning reaches a normalised score of within episodes and transfers without retraining to unseen initial states (–), where a recurrent baseline drops to –. A DQN that observes the latent eventually matches or exceeds it, consistent with the certificate's slack.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.