acceptodds
Under review as a conference paper at ICLR 2027

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss

Abstract

Many modern applications of deep learning involve training a neural network via a one-step prediction loss (e.g., regression, cross-entropy), but deploy the network by rolling out along its own predictions. Key examples include autoregressive language modeling, flow-based generative modeling, and robot policy learning. It is well-documented that these settings induce a phenomenon we call **test-time feedback** (TTF): the mismatch between the training/validation loss and downstream metrics of interest, such as task success rate and generation quality, which grows with task length. While data curation, architecture, and objective design have been proposed to combat train-test shift in TTF settings, this paper proposes *optimization* as a new design axis to mitigate error accumulation. Specifically, we introduce a new optimization paradigm called **double-preconditioning** () uniquely tailored to the challenges of TTF. combines gradient-wise preconditioning, as in and , with activation-wise preconditioning (), such as in . We show that the addition of yields a *drop-in intervention* for increasing downstream model performance across a range of TTF settings. Interestingly, these gains in test-time performance *do not consistently accompany improvements in validation loss*, opening new questions about how to properly evaluate models trained with one-step supervised objectives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.