Learning Where to Constrain: Meta-Learned Trust Allocation for Language-Model Reasoning
Abstract
Policy updates in auto-regressive reasoning change both token probabilities and the prefixes on which later decisions depend. Recent trust-region methods for LLM reinforcement learning account for token position and accumulated prefix divergence to improve training stability. However, a fixed or prespecified allocation may not suit different response lengths or changing task distributions. To study whether these constraints should instead be adaptive, we introduce Meta-TA (Meta-learned Trust Allocation), which learns an admissible position-weight floor and prefix-divergence budget by differentiating through the policy update. We show why a one-step update starting from the rollout policy provides no allocator gradient through an inactive constraint. A detached warm-up followed by a differentiable soft-constrained update enables allocator learning when the constraints become active. We then freeze the learned allocation and evaluate its transfer to fresh policy training. Our main evaluation uses the Qwen model family on mathematical reasoning, and we extend the study to abrupt task changes, gradual distribution shifts, and recurring tasks to examine adaptation and retention. To mitigate allocator overfitting, we select among saved allocator checkpoints using a separate subset and examine whether repeated updates on the query subset lead to overfitting. Decoupling allocator-regularization from policy-entropy helps distinguish restrictions on the learned update rule from changes in policy exploration. Together, these comparisons examine when learned constraints improve LLM reinforcement learning reasoning and if those gains justify additional cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.