acceptodds
Under review as a conference paper at ICLR 2027

Test-Time RL for Diffusion Language Models: Geometry-Guided Selective Updates

Abstract

Test-time reinforcement learning (TTRL) enables models to adapt from unlabeled test-time experience, but agreement among sampled answers alone does not establish which targets are reliable enough to reinforce. We introduce Geo.-LiM, a framework for persistent, label-free adaptation of diffusion language models (dLLMs) that exploits complementary evidence from NFE-matched denoising geometries. Geo.-LiM selectively triggers updates by combining strict pooled-majority support across geometries with novelty relative to a deterministic frozen-reference answer, and reinforces only trajectories supporting eligible candidates under fixed-reference regularization. Across three dLLMs and five benchmarks, Geo.-LiM improves post-adaptation performance over the unadapted backbone in 27 of 30 model-task-length settings. Under the same eight-rollout evidence budget, it also outperforms our consensus-based TTRL baseline in 23 of 30 settings while admitting fewer prompts and requiring fewer optimizer steps in every setting, reducing their aggregate counts by 74.2% and 64.4%, respectively. These results show that denoising geometry provides useful evidence for identifying informative test-time updates, enabling selective and efficient persistent adaptation of dLLMs without target labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.