TTS-RiskArena: Unsafe at Scale? Evaluating and Mitigating Risk in Test-Time Scaling
Abstract
Test-time scaling (TTS) improves the capability of large language models by allocating additional inference-time computation to explore multiple reasoning trajectories. However, its safety implications remain largely unexplored. A model that appears aligned under standard decoding may behave differently when given larger search budgets, more samples, or iterative refinement, raising a fundamental question: Does TTS preserve safety, or can it amplify unsafe behavior? This work introduces TTS-RiskArena, a benchmark to systematically evaluate safety under test-time scaling. TTS-RiskArena spans seven safety-critical domains, including health, finance, coding, child, scientific discovery, travel, and biochemistry, and enables controlled evaluation across diverse TTS methods and inference budgets. The benchmark further organizes prompts across multiple abstraction levels, making it possible to analyze how prompt specificity interacts with safety under increasingly exploratory inference. Further, we argue that TTS is fundamentally a search over latent reasoning space, with the decoding process verbalizing the trajectories it reaches. Consequently, the chance of producing a safe response is largely determined by whether early inference makes safe regions of that space more reachable. Based on this intuition, we propose a contrastive decoding method that guides test-time exploration toward safer regions of the output space. Experiments show that the proposed method consistently improves safety while retaining the benefits of TTS.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.