acceptodds
Under review as a conference paper at ICLR 2027

Scalable and Efficient Learning Without Rewards in High Dimensional MDPs

Abstract

From its theoretical foundations to mathematical and scientific reasoning of frontier reasoning models, reinforcement learning has been the cornerstone of striking progress in learning complex policies directly from high-dimensional representations. Nonetheless, a significant limitation of reinforcement learning remains its reliance on a predefined, extrinsic reward signal, a foundational assumption that often fails when the computational cost of building a reward function exceeds that of learning the policy itself. In this paper, we investigate the problem of sequential decision-making in high-dimensional MDPs where reward information is absent. We introduce harmonic learning, a theoretically well-founded paradigm for learning in high dimensional MDPs where the explicit reward is not provided. By leveraging the uncertainty principle of harmonic analysis, our analysis and method provide a theoretical basis for effective and efficient learning with accelerated training. The theoretical and empirical analysis reported in our paper demonstrates that our method achieves substantially higher performance with significant efficiency, resulting in stable and resilient policies that can generalize to uncertain environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.