acceptodds
Under review as a conference paper at ICLR 2027

Tasks Do Not Tensorize: Learning When to Reset or Continue

Abstract

Meta-reinforcement learning usually treats the grouping of source trajectories into tasks as fixed. We study an acquisition model in which the learner decides whether the next controlled trajectory should retain the current hidden task or begin an independent one. This changes the statistical experiment. For every , we construct finite meta-MDPs in which every episodewise experiment has zero information about two target-incompatible models while one bounded persistent experiment is informative; under perturbation the persistent-to-episodewise rate ratio is arbitrarily large. For every fixed , the zero-terminal-value -step receding-horizon information optimizer can likewise be arbitrarily inefficient. For finite persistent meta-MDPs with bounded task depth, a finite observational quotient, and disjoint single-answer cells, we characterize the exact fixed-confidence source cost over all randomized history-dependent experiment trees. The characteristic information is a Bellman–Chernoff game: a dynamic program chooses the next episodic policy and whether to stop, while a least-favorable target alternative prices information. Certified column generation computes a sparse design, and vanishing certified optimization error preserves the characteristic constant. A cross-modal calibration model illustrates how continuation reveals correspondence while resetting supplies independent breadth. Thus task grouping is not bookkeeping: it is an information-control action.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.