acceptodds
Under review as a conference paper at ICLR 2027

TEACH ME NO REGRET: ACCESSIBILITY AND REALIZABILITY IN REGRET-BASED POST-TRAINING

Abstract

Regret-minimization fine-tuning (RMFT) trains a decision policy on selected low-regret trajectories. Its success depends both on generating useful trajectories and on reproducing their behavior with the information available at deployment. We develop a scenario-wise analysis that separates these two requirements. For regret distributions with arbitrary atoms and ties, we derive a finite-iteration bound combining initial lower-tail accessibility with a selection-to-deployment discrepancy. A finite-sample result separates trajectory sampling, score error, and policy fitting, while margin conditions translate regret into rollout-averaged action error. We also solve repeated population imitation in a binary model with noisy observations: the policy converges geometrically to the best admissible decision rule, although an information-dependent gap to the evaluator's oracle persists. Experiments with a lightweight language model in stationary, contextual, and semisynthetic recommendation bandits show contrasting responses to direct RMFT and teacher-distilled initialization. The resulting framework distinguishes the quality of selected trajectories from the performance achieved by the deployed policy and provides a basis for interpreting initialization-dependent post-training gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.