acceptodds
Under review as a conference paper at ICLR 2027

DIVTHINK: Parallel Thinking via Diversified Chain-of-Thought Sampling

Abstract

Test-time reasoning has become a central paradigm for scaling Large Language Models (LLMs), but parallel Chain-of-Thought (CoT) sampling often wastes computation on redundant trajectories that repeat similar intermediate steps and converge to the same final answer. Existing alternatives either refine trajectories sequentially, as in MCMC power sampling, or require costly post-training, as in GRPO. To bridge this gap, we propose DIVTHINK, a training-free decoding method that coordinates parallel CoT trajectories through lightweight repulsive logit adjustments. At each step, DIVTHINK downweights tokens already preferred by peer trajectories, encouraging complementary reasoning paths while preserving the base model's local preferences and the throughput advantage of parallel generation. Across mathematical reasoning, code generation, and scientific question-answering benchmarks, DIVTHINK improves accuracy by 7.6 points over Base@8 and 4.0 points over Low-temperature@8 on average, with gains up to 22 points, while using far fewer generated tokens than MCMC Power Sampling and avoiding the rollout and optimization cost of GRPO-style post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.