CATER: Competence-Adaptive Training for Efficient Reasoning
Abstract
Efficient reasoning requires models to spend computation where it remains necessary. Yet long Chain-of-Thought (CoT) models can remain verbose on problems they already solve reliably, while uniform pressure to shorten responses can compromise reasoning on problems they have not yet mastered. Existing reinforcement learning approaches promote efficiency through outcome-based rewards, but sparse feedback provides limited guidance on how to eliminate unnecessary reasoning without sacrificing correctness. To resolve this problem, we propose CATER (Competence-Adaptive Training for Efficient Reasoning), a difficulty-adaptive framework that combines reinforcement learning with on-policy self-distillation, using the student’s online success rate to assign a training signal to each problem. For consistently solved problems, concise distillation preserves existing capabilities with fewer reasoning tokens. For inconsistently solved problems, correctness-based reinforcement learning reinforces successful reasoning strategies. By adapting supervision to the student’s evolving competence, our framework advances reasoning compression and capability development together. On Qwen3-4B, it reduces reasoning tokens by 39.1% on average across five benchmarks while maintaining comparable overall accuracy. These efficiency gains extend beyond mathematics to scientific reasoning and code generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.