acceptodds
Under review as a conference paper at ICLR 2027

CATER: Competence-Adaptive Training for Efficient Reasoning

Abstract

Efficient reasoning requires models to spend computation where it remains necessary. Yet long Chain-of-Thought (CoT) models can remain verbose on problems they already solve reliably, while uniform pressure to shorten responses can compromise reasoning on problems they have not yet mastered. Existing reinforcement learning approaches promote efficiency through outcome-based rewards, but sparse feedback provides limited guidance on how to eliminate unnecessary reasoning without sacrificing correctness. To resolve this problem, we propose CATER (Competence-Adaptive Training for Efficient Reasoning), a difficulty-adaptive framework that combines reinforcement learning with on-policy self-distillation, using the student’s online success rate to assign a training signal to each problem. For consistently solved problems, concise distillation preserves existing capabilities with fewer reasoning tokens. For inconsistently solved problems, correctness-based reinforcement learning reinforces successful reasoning strategies. By adapting supervision to the student’s evolving competence, our framework advances reasoning compression and capability development together. On Qwen3-4B, it reduces reasoning tokens by 39.1% on average across five benchmarks while maintaining comparable overall accuracy. These efficiency gains extend beyond mathematics to scientific reasoning and code generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.