acceptodds
Under review as a conference paper at ICLR 2027

Spatial But Not Temporal Reward Sparsity Drives Slow Policy Learning

Abstract

A primary difference between reinforcement and supervised learning is that rewards are generally sparse in both space and time: few actions yield high reward (spatial sparsity), and rewards may only be provided after temporally extended sequences of actions (temporal sparsity). How these factors impact policy learning dynamics is undertheorized, especially because it is difficult to design a theoretical setting interesting enough to capture relevant phenomenology, but simple enough to treat analytically. In this paper, we propose and analytically solve a minimal model of learning to pursue a moving target which naturally incorporates both spatial and temporal reward sparsity. We show mathematically that spatial sparsity can produce exponential slowdowns in learning, and that temporal sparsity is comparatively benign, only yielding a comparatively short additional phase of pursuit trajectory refinement. We also show that the presence of highly variable environmental state features can—without complex function approximators—produce an implicit curriculum effect which makes learning from sparse rewards polynomially rather than exponentially slow. We provide a wealth of relevant calculations, whose utility is in part that they sharpen our intuitions about reward sparsity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.