acceptodds
Under review as a conference paper at ICLR 2027

DARS: Dynamic Alignment Reward Steering for Robust Adaptive Reasoning

Abstract

Adaptive reasoning models allocate inference-time computation to tasks of varying difficulty, enabling an efficiency–performance trade-off. Yet two key flaws exist: under-specified rewards that fail to penalize misaligned reasoning budgets, and an “elastic” failure mode in which models tend to revert to non-adaptive reasoning. To this end, we propose DARS (Dynamic Alignment Reward Steering), a novel Reinforcement Learning Finetuning (RLF) framework that combines two synergistic strategies—Minimal Sufficient Reasoning (MSR)-guided optimization and Fisher regularization—to improve the robustness of adaptive reasoning fine-tuning. Concretely, MSR-guided optimization begins with an oracle-calibrated difficulty estimator that assigns each instance to a difficulty tier. Building on this, MSR-driven reward explicitly penalizes both short “arrogant” reasoning and prolonged over-thinking while promoting diverse reasoning paths. To stabilize RLF and mitigate elastic reversion, we further incorporate a Fisher-informed regularizer that constrains updates to high-importance parameters according to the Fisher information matrix. Experiments demonstrate that DARS prevents catastrophic collapse and achieves robust adaptivity across diverse benchmarks. On standard complex reasoning tasks like MATH, it substantially improves the accuracy–efficiency trade-off (+0.6% accuracy with 3.5% fewer tokens on 14B), and on extremely challenging benchmarks like AIME, DARS maintains a much narrower token fluctuation range under adversarial attacks (e.g., reducing the token changes span by over 30% on 14B compared to strong baselines), and effectively resists severe mode reversion.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.