Thanks for Failing: Better Test-Time Search from Wrong Attempts
Abstract
Test-time scaling (TTS) methods are central to both LLM evaluation, and post-training pipelines. They not only help improve inference-time accuracy, but also TTS-inspired methods support rejection sampling and search- or verifier-guided trajectory generation for reinforcement learning. For rollout generation, TTS methods can often be decomposed into two primitives: independent sampling from a shared context, and sequential sampling conditioned on previous attempts. Recent work finds that independent sampling often outperforms sequential sampling, yet we observe that these two primitives have competing strengths. Independent sampling explores more diverse solutions, but cannot leverage prior failures. Sequential revision depends on prior failures, but becomes less diverse and is often trapped by failed reasoning. In this paper, we systematically study four methods to combine the predictions of independent and sequential sampling strategies. Through this analysis, we propose Failure Logit Extrapolation (FLEx), a simple, training-free, contrastive decoding algorithm that extrapolates the difference between token-level logits generated independently and sequentially. Across six math, coding, and instruction-following benchmarks, and eight models ranging from 4B to 235B parameters, FLEx improves pass@8 over independent sampling by 3.3% and 11.5% over sequential sampling on average. On the four mathematics benchmarks, it also increases answer diversity, measured by the fraction of unique answers, by 5.9%; on problems that all three methods fail to solve, this diversity advantage grows to 7.3%. Its gains consistently increase as more failed attempts are retained in its context and as more sampling budget is provided. Finally, we show that FLEx directly benefits post-training pipelines. When used for rejection sampling, FLEx-collected trajectories result in higher downstream accuracy than independently sampled trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.