acceptodds
Under review as a conference paper at ICLR 2027

First Generate, Then Delete: Test-Time Thinking-Token Deletion for Multimodal Reasoning

Abstract

Large reasoning models typically solve complex problems by generating long chain-of-thought traces before producing final answers. In this work, we observe a phenomenon that not every generated thinking token is equally beneficial for subsequent answer decoding. Motivated by this observation, we propose Test-time Token Deletion (TTD), a first-generate-then-delete framework for test-time reasoning intervention. Given a question, the model first produces an initial reasoning trace, and TTD then selectively deletes a subset of thinking tokens and re-decodes the final solution conditioned on the remaining trace. To select tokens for deletion, we instantiate TTD with a training-free entropy-guided rule and a test-time mask-learning variant optimized with forward-only evaluations, while keeping the base model frozen. Extensive experiments on various multimodal reasoning benchmarks show that TTD improves reasoning accuracy while reducing response length. Our findings suggest that selective thinking-token pruning provides a complementary axis of test-time scaling beyond generating more tokens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.