acceptodds
Under review as a conference paper at ICLR 2027

Response-Conditioned Optimizer Routing On One Irreversible Trajectory

Abstract

Many stochastic optimizers are available for training a model, and it is rarely clear in advance which one to use. Worst-case convergence bounds can be a pessimistic guide; A/B tests need one training run per candidate, which, for a single large model, can cost more than a good choice saves; tests on smaller models can be misleading. We therefore study how to choose the optimizer while the model itself is being trained. Such a run has only one model state: updates cannot be undone, and optimizers that were not chosen cannot be tried from the same state. At each step, a stochastic gradient is revealed first, and a router then chooses which candidate algorithm uses it to produce a new testing point. We compare the expected terminal loss with that of the best fixed optimizer run from the same initial state with the same budget of \(B\) gradients and updates, without resets, model copies, or extra queries. We propose three methods that share one model of the objective, fitted during a persistently exciting calibration phase that explores all decision-relevant directions. PE-Commit chooses before the first post-calibration gradient, PE-Post chooses once after seeing it, and PE-RCR may switch again at later decision points. For unknown strongly convex quadratics and optimizers with contractive affine-linear state updates, we prove explicit finite-sample bounds on the positive part of the expected pseudo-regret of all three methods. The bounds hold in expectation and assume bounded, conditionally mean-zero noise with a fixed covariance, which is stronger than generic minibatch noise; PE-RCR pays an estimated term at every decision point. An exact two-step example shows that scoring by the remaining horizon can be the unique optimal gradient-contingent plan and can strictly beat one-step greedy selection that also sees the gradient. A lower bound shows that no router can identify the better next method uniformly without information along decision-relevant directions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.