acceptodds
Under review as a conference paper at ICLR 2027

Learning Controlled Dynamics through Obedient Recommendations

Abstract

Learning controlled dynamics for model-based reinforcement learning requires informative actions, which a learner may need to induce strategic receivers to execute voluntarily. We study how recommendation visibility changes which experiments can be implemented and which dynamics can be identified. We construct a transfer-free two-stage system with informed, far-sighted receivers, holding dynamics, utilities, sender objectives, and learner feedback fixed across public and receiver-private recommendations. The architectures have equal worst-equilibrium known-model values, yet minimax regret under worst-case sequential-equilibrium selection is linear publicly and logarithmic privately for fixed model alternatives. Private signals uniquely implement actions that reach and activate an informative transition, whereas public signaling admits a common uninformative equilibrium selection. For general mechanism classes under prescribed obedience, we introduce the obedient observability closure, which iterates model elimination and the resulting expansion of obedient experiments. With finite model classes, compact mechanism classes, and exact diagnostic and oracle access, sublinear regret with all-round obedience at vanishing failure probability is possible exactly when the models in each terminal closure share an optimal obedient mechanism. Information architecture thus governs learnability by determining which physical experiments can be strategically implemented.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.