ASTra: Mitigating overthinking via tracking attention sink
Abstract
The overthinking problem has recently attracted significant attention from researchers. While the introduction of chain of thought substantially enhances the model's ability to solve complex tasks, the overthinking problem emerges when solving relatively simple questions. In these cases, overthinking can be observed as repeated verifications and divergent thoughts, which do not translate into improved problem-solving, yet generate considerable computational overhead. In this paper, we propose ASTra (Attention Sink Tracker), an inference-time method to mitigate overthinking problems in LLM. By monitoring attention weight distribution, we discovered an optimal truncation point to prompt models to terminate the chain of thought. Extensive experiments show that our method reduces 15% to 70% of chain of thought tokens, while maintaining final output quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.