acceptodds
Under review as a conference paper at ICLR 2027

Gate-Override: Cost Amplification Attacks on Adaptive Reasoning Models

Abstract

Adaptive reasoning models save computation by choosing when to answer directly and when to reason at length. This choice also gives an attacker a way to increase computation on questions the model already answers correctly. We present Gate-Override, an audit that tests this risk by linking each request’s executed reasoning mode, answer correctness, and generation cost. Under white-box access, we optimize a shared 10-token suffix and insert its token IDs before the assistant-generation marker. Across five model–task settings spanning 1.5B–15B parameters, the resulting attacks change reasoning modes and increase the cost of correct answers. On AdaptThink-1.5B/GSM8K, the suffix switches 75 of 84 direct-answer requests to reasoning. On Phi-4-RV-15B/MMLU-Pro, it switches 93 of 100, versus zero for either random suffix; 48 of 52 complete, originally correct pairs remain correct while switching routes and costing at least 1.5 times as much. Our paired metrics reveal cost differences hidden by equal route-switch rates and accuracy. Forcing the same input through each route attributes most extra computation to the route change. Re-optimizing against a hardened route head restores substantial attack success, and matched serving experiments show added latency for clean requests. These findings make reasoning-mode selection a concrete target for resource-security evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.