acceptodds
Under review as a conference paper at ICLR 2027

The Price of Cross-Action Sharing: Regret Bounds and Information Limits in RL

Abstract

Observed environmental inputs can support cross-action reuse in reinforcement learning (RL), but an input sample alone does not reveal unexecuted outcomes when conditional responses are unknown. Which observations can then be pooled, and what does validating their reuse cost? We study finite episodic Markov decision processes with unknown input-sharing groups, conditional transition kernels, and reward means. A response-mediated transfer characterization distinguishes equal input laws from equal induced transitions and quantifies the bilinear interaction between input and response estimation errors. Certified Response-First Optimism (CRFO) separates group certification from estimation, retains individual estimates, and bounds conditional continuation values before averaging input laws. This ordering charges response uncertainty to executed transitions. Its separation-dependent, high-probability regret bound is at most twice the smaller of the complete individual-learning and certification-plus-pooling bounds, including identification, response, and reward costs. Targetwise confidence intersection can tighten jointly valid upper estimates without changing the proved regret order. A streaming temporal-difference extension reduces conditional-response storage through a distinct historical-target analysis. For two-policy selection with known reward formulas, matching information bounds reveal a sharper distinction. At a shared reference, let scale cancellation, be the common successor-mean gap, and the response contrast, so the policy gap is . Under uniform correctness over models admitting unequal input laws without positive separation, the pointwise optimal expected episode count is , versus when input-law equality is known, where is the error probability. A product-aware stopping rule attains these orders with reference-tuned initialization and adapts to unknown resolution with an explicit iterated-logarithmic overhead. At a fixed successor gap, vanishing response contrast changes the small-cancellation cost from inverse-square to inverse-first-power because the sharing-induced alternative requires joint input–response perturbations. Thus, neither group count nor policy-gap size alone determines sharing cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.