Parallel Test-Time Scaling with Multi-Sequence Verifiers
Abstract
Parallel test-time scaling—generating multiple candidate solutions for a single problem—is a powerful technique for improving large language model performance, but it faces two related challenges: selecting the correct solution from a large candidate pool is difficult, and generating every candidate to completion can be slow. Predicting candidate correctness can address both problems by guiding answer selection and informing when to stop generation. However, verifiers that score candidates independently cannot exploit the relationships between candidates' underlying representations. To address this, we introduce the Multi-Sequence Verifier (MSV), a lightweight verifier that predicts each candidate's correctness by jointly attending over the hidden representations of all sampled candidates, rather than processing each one in isolation. We apply MSV's predictions to best-of-N selection, confidence estimation, and parallel early stopping, evaluating each use case separately. Across challenging mathematical reasoning benchmarks, MSV improves best-of-64 accuracy by up to 6% relative to the strongest baselines, and a streaming variant of MSV improves the estimated accuracy-latency tradeoff over the evaluated baseline verifiers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.