A Conservation Law for Learning in Strategic Systems
Abstract
Standard exploration diagnostics measure how much a policy moves, but not whether that movement remains informative after the data-generating process responds. We introduce strategic observability to measure how strongly policy motion separates latent environments after this response, and define strategic capacity as motion times squared observability. For binary environments, mutual information is bounded above by expected capacity and below by a truncated almost-sure capacity floor. Under local Fisher geometry, the quadratic exploration cost decomposes orthogonally into identifiable and null components. Their relative weights are the squared cosine and sine of the angle between policy motion and the identifiable subspace; the identifiable component controls learning up to the conditioning of the identification operator. Local characterisation results further select Fisher separation energy as the learning currency and the symmetric Fisher quadratic form as the cost. The finite-step quantities used in practice are measurable surrogates for, rather than exact copies of, these geometric objects. Empirically, six policies with exactly equal budget, average separation energy, switching, exposure, and action entropy separate only through minimum pairwise separation: the two with zero minimum perform at chance, whereas the remaining four achieve error below . A language-model intervention likewise shows that greater motion along an uninformative direction can yield negligible separation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.