acceptodds
Under review as a conference paper at ICLR 2027

Robust Distributed Test-Time Scaling

Abstract

Test-time scaling (TTS) improves the reasoning performance of large language models by allocating more computation at inference time. At scale, the independent generations underlying TTS can naturally be distributed across multiple workers and aggregated into a final prediction. This paradigm, however, inherits a fundamental challenge from distributed computing: tolerating faulty or compromised workers. To address this challenge, we develop proof-of-thought consensus, a robust aggregation framework for distributed TTS. The key insight is that producing a chain-of-thought (CoT) trace that remains statistically consistent with the model’s decoding distribution is inherently costly for an adversary. We then introduce an efficient CoT verification mechanism that turns this asymmetry into a practical filter by detecting subtle distributional inconsistencies in manipulated CoT traces before aggregation. We complement the design with theoretical and adversarial analyses: we prove high-probability acceptance guarantees for honest submissions, and formalize three strong attack strategies for stress-testing. Experiments show that, even with up to 30% malicious participants, we limit accuracy degradation to 6.7% on average relative to an oracle baseline that aggregates only the clean samples, while incurring only a 1% additional inference-time overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.