acceptodds
Under review as a conference paper at ICLR 2027

TIDES: Test-time Inference Drift Exploitation via Scaling

Abstract

We propose TIDES, a reasoning attack method that exposes a previously unrecognized failure mode of test-time scaling: as reasoning traces lengthen, model performance can degrade sharply rather than improve. Unlike prior attacks on large reasoning models (LRMs), TIDES exploits computation depth under test-time scaling as an implicit activation condition, producing degradations that remain difficult to detect under short-budget evaluations. Methodologically, we define a Depth-Guided Latent Tracker (DLT), a depth-based tracker that injects microscopic steering vectors into intermediate reasoning representations and combines them with on-policy distillation to embed computation-dependent behavior in LRMs. Theoretically, we model latent space as a depth-indexed dynamic process and derive a worst-case propagation bound characterizing the potential sensitivity of long reasoning trajectories to small intermediate perturbations. Empirically, we evaluate TIDES on multiple reasoning benchmarks using three strong reasoning models, Phi-4-Reasoning, DeepSeek-R1-Distill-Qwen-7B, and DeepSeek-R1-Distill-Llama-8B, where it outperforms strong reasoning attack methods such as DecepChain and BadChain. Notably, TIDES achieves an average 26.8% relative improvement in RAS over the strongest baseline while largely preserving utility under short reasoning budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.