acceptodds
Under review as a conference paper at ICLR 2027

Exploration-Free Learning Needs an Exact Plug-In: Information Floors and Sharp Regret Constants for Event-Triggered Estimation

Abstract

Exploration is often viewed as the price of learning to act: a policy that simply plugs in its current parameter estimates can fail to gather the information it needs, and in online control of unknown linear systems minimax regret grows as sqrt(T). We study problems with an information floor, meaning that information about the unknown parameters arrives at a rate that no admissible policy can drive to zero, and show that this condition alone is not enough. Our main result is a dichotomy that is invisible in continuous time. Write r for the residual of the learner's plug-in map at the true parameter, i.e. the gap between the threshold it returns there and the optimum of the sampled problem. If r = 0, certainty equivalence attains regret C_Δ log τ + O(1) with an explicit constant; if r ≠ 0, its regret is linear at rate ½ J″_Δ(β*_Δ) r². The natural continuity-corrected map has r = Θ(Δ) and is therefore linearly regretful, so a statistically consistent estimator does not suffice: the map it is plugged into must be consistent for the sampled comparator too. A van Trees bound over shrinking neighborhoods matches C_Δ exactly for every locally consistent policy, one whose regret is o(τ) uniformly on a neighborhood of the truth. So within that class C_Δ is the price of parameter uncertainty rather than an artefact of one algorithm. We also sharpen the underlying discretization theory: the optimal sampled threshold is β*_Δ = β* − 0.5826 σ√Δ − σ²Δ/(8β*) + o(Δ), while the optimal cost rises by exactly λσ²Δ/6 + o(Δ). We instantiate the framework twice, through event-triggered remote estimation over a lossy channel with unknown volatility and loss rate and through repeated sequential testing with unknown noise scale, and the same dichotomy appears in both. All cost functionals come from a deterministic Fredholm solve validated against Monte Carlo: across eight parameter settings the measured regret slope matches C_Δ with ratio 1.00 ± 0.11, and the predicted linear rates of inexact plug-ins are confirmed to within 5%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.