When Reasoning Fails to Stop: Prompt-Induced Resource Exhaustion in Black-Box Reasoning Models
Abstract
Reasoning large language models allocate additional inference-time computation to improve multi-step reasoning and complex problem solving. Existing studies have demonstrated excessive computation and high resource consumption by constructing or optimizing inputs. However, attack effectiveness alone does not explain the task conditions under which resource exhaustion occurs. We therefore investigate which prompt conditions induce continued computation and resource exhaustion in a black-box setting limited to normal question-answering requests. We construct Reasoning-Bench with five task types and three task-complexity levels, evaluating 400 prompts on 10 reasoning models to obtain 4,000 runs. Resource statistics, execution-behavior analysis, and condition ablations yield six findings: task complexity cannot directly determine resource consumption; code, cryptography, and algorithmic-reasoning tasks tend to exhibit higher resource consumption; models tend to execute first and judge later when tasks are declared solvable; continuously decomposable tasks encourage local exploration; missing or unverifiable information encourages multiple hypotheses and continued reasoning; and detailed-process requirements further increase resource consumption. To test whether these findings guide resource-exhaustion task construction, we propose ReasoningSink, a resource-exhaustion prompt generation framework for black-box reasoning models. The framework defines prompt conditions and model-behavior checks through State Constraints, organizes tasks through Dynamic Retrieval-Augmented Generation, and revises candidates through Execution-Feedback Optimization. Across the 10 reasoning models, ReasoningSink produces averages of 35.4K Completion Tokens (CT) and output-to-input Amplification Ratio (Amp). For every evaluated model, ReasoningSink achieves higher CT and higher Maximum-Budget Hits (MaxH) than the comparison methods. Reasoning text from ReasoningSink-generated tasks again exhibits executing before judging, local exploration, hypothesis-space expansion, and detailed-process elaboration, validating the findings' guidance for resource-exhaustion task construction. Code, dataset, and evaluation scripts are available at the anonymous repository Reasoning-Bench-and-ReasoningSink.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.