ReCAT: Recurrent Composable Activation Transition for Test-Time Scaling of Diffusion Models
Abstract
Diffusion test-time scaling typically spends additional compute on generating, evaluating, or revising multiple candidates. We introduce ReCAT, a sequential scaling method that instead refines a single evolving internal state within one sampling trajectory. At a fixed intervention point, the frozen generator computes an intermediate activation once; a lightweight prompt- and state-conditioned controller then recurrently applies the same learned transition for a number of steps chosen at inference before generation resumes. Each additional refinement step requires only one controller call and no online reward queries. Reward-Ascent Matching (RAM) learns this transition from black-box rewards by fitting reward changes to activation displacements within sample groups and distilling the resulting secant directions across prompts. The resulting controller can be used directly; an optional second stage, Reward-Flow Calibration (RFC), further calibrates it on states reached by the controller's own updates using antithetic continuation probes. Neither stage backpropagates through the generator or reward model. Across three diffusion backbones and four target rewards, RAM alone yields consistent reward improvements, RFC provides additional gains, and reward gains increase with the number of refinement steps throughout the calibrated range. ReCAT also composes with search-based test-time scaling methods, improving Best-of-N search and gradient-based noise search.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.