Scalable and Efficient Learning Without Rewards in High Dimensional MDPs
Abstract
From its theoretical foundations to mathematical and scientific reasoning of frontier reasoning models, reinforcement learning has been the cornerstone of striking progress in learning complex policies directly from high-dimensional representations. Nonetheless, a significant limitation of reinforcement learning remains its reliance on a predefined, extrinsic reward signal, a foundational assumption that often fails when the computational cost of building a reward function exceeds that of learning the policy itself. In this paper, we investigate the problem of sequential decision-making in high-dimensional MDPs where reward information is absent. We introduce harmonic learning, a theoretically well-founded paradigm for learning in high dimensional MDPs where the explicit reward is not provided. By leveraging the uncertainty principle of harmonic analysis, our analysis and method provide a theoretical basis for effective and efficient learning with accelerated training. The theoretical and empirical analysis reported in our paper demonstrates that our method achieves substantially higher performance with significant efficiency, resulting in stable and resilient policies that can generalize to uncertain environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.