acceptodds
Under review as a conference paper at ICLR 2027

Learning While Generating in a Single Forward Pass

Abstract

Deployed LLMs constantly receive feedback from users, verifiers, and task outcomes. Test-time adaptation uses this feedback so that each interaction improves the next one. Today, this learning needs a second pass after feedback arrives, which either reruns the model over the response or keeps all of its activations until then. We observe that the model computation in this second pass does not depend on the feedback, so it can be done in the forward pass that generates the response. Based on this observation, we propose **One-Pass Forward-Trace Adaptation (OFTA)**, which learns while generating in a single forward pass. As the model generates each token, OFTA also records how the log-probability of the response changes along a random parameter direction. This design brings three benefits. First, it removes the second pass. When feedback arrives, OFTA turns the record into an update with a few simple operations, without running the adapting model again, using backpropagation, or keeping activations. Second, the record works for immediate and delayed feedback, and one record supports many feedback losses. Third, moving the computation into generation costs nothing in update quality. Theoretically, OFTA gives the same directional derivative as a separate pass after feedback, so it keeps the unbiased estimate of forward-gradient methods. Experiments on Qwen3 series models show that, in the tested settings and counting all extra work, OFTA lowers request time, applies updates almost as soon as feedback arrives, uses much less memory than backpropagation, and learns as well as a separate pass. In simpler terms, OFTA lets a model learn from a response in the same pass that writes it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.