acceptodds
Under review as a conference paper at ICLR 2027

PRICE: Learning What to Recompute for Reasoning in Diffusion Language Models

Abstract

Diffusion language models can revise outputs, but choosing what to recompute is difficult. Local repair may leave inconsistent downstream results. Recomputing all descendants may waste useful work. We propose PRICE (Predictive Repair via Intervention-trained Candidate Evaluation). It chooses semantic repair scopes using learned estimates of final reward and remaining cost. A predicted dependency graph proposes a bounded pool of executable actions. Offline continuations from the same partial state train estimates of reward and cost differences. Inference runs one action without trying alternative continuations. The executor hides predicted affected text from the model. Cached values return only after a check without reference answers; its cost is counted. Our analysis shows when selective repair beats full closure and when learned selection preserves this gain. It bounds regret from missing candidates and value errors. Across three instruction-tuned diffusion backbones and four reasoning benchmarks, evaluated settings improve scores over full closure with fewer model evaluations. Gains depend on candidate coverage, value accuracy, and the reliability and full cost of checked reuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.