Inter-Trajectory Dependence Limits Answer Coverage in Local dLLM Repair
Abstract
Test-time compute scaling improves large language model performance by allocating more computation during inference. For diffusion large language models (dLLMs), sequential repair offers a natural instantiation: it iteratively remasks low-confidence tokens and regenerates them while retaining the surrounding completion. However, additional repair rounds often yield limited performance gains. While recent work has studied inference dynamics within individual trajectories, our focus is inter-trajectory dependence: whether successive repair candidates explore distinct answer states. We introduce a diagnostic framework that combines answer-level metrics, hyperparameter sweeps, and remasking interventions to characterize this dependence. We assess pairwise answer agreement, unique-answer counts, oracle pass@, and correction/regression rates across four dLLMs and three verifiable tasks. We find that sequential local repair produces highly correlated candidates with substantially lower answer coverage than independent sampling. The evaluated local-remasking variants do not reliably close this gap. These findings suggest prioritizing trajectory-level diversification when allocating test-time compute for answer coverage, with local repair used for refinement. Our study adds an inter-trajectory perspective to the analysis of dLLM inference dynamics and provides practical guidance for allocating inference compute.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.