Enhancing Parallel Tempering for Test-Time Steering in Masked Diffusion Models
Abstract
Test-time steering aims to sample from reward-tilted distributions defined by a pre-trained model and an external reward function, without updating the model's parameters. However, samplers may struggle to move between high-reward regions separated by low-probability barriers, leaving important modes unexplored. Sequential Monte Carlo may concentrate particles around early-discovered modes, while replica exchange across diffusion noise levels may mix slowly when reward-agnostic exchange proposals are frequently rejected. We propose Temperature-Enhanced Masked Proposal and Exchange for Reward Steering (TEMPERS), a reward-centered parallel tempering framework for masked diffusion models that operates on complete sequences. Remask-and-denoise proposals enable reward-based acceptance at every local update, allowing each replica to refine candidates at its reward temperature before exchange. In an analytic example with a known reward-tilted distribution, TEMPERS closely approximates the target region probabilities. Experiments on protein and DNA sequence design and sentiment-controlled text generation further show higher rewards than the evaluated test-time steering baselines, alongside competitive output quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.