acceptodds
Under review as a conference paper at ICLR 2027

Efficient LLM Reasoning via Answer-Guided Posterior-to-Prior Transfer

Abstract

Although large language models achieve strong reasoning capabilities through reinforcement learning, they often generate unnecessarily long reasoning trajectories, leading to substantial inference costs. Conventional length-based reward engineering curtails redundancy but overlooks the fundamental sampling bottleneck: correct-yet-concise trajectories remain severely sparse under question-only exploration, hindering further gains in reasoning efficiency. Inspired by cognitive science, where knowledge of outcomes suppresses ineffective trial-and-error, we leverage training-time reference answers as privileged signals to guide exploration. We theoretically prove that, under a non-negative covariance condition between correctness likelihood and comprehensive utility, the ideal Bayesian answer-conditioned posterior yields no lower expected utility than the question-only prior, with strict improvement when the covariance is positive, providing a principled foundation for answer-guided exploration. To bridge the posterior-prior gap and prevent the transfer of privileged shortcuts, we propose VPG-EA (Variational Posterior Guidance and Efficiency Awareness). Motivated by a variational lower-bound perspective, VPG-EA uses Cross-View Validation to select trajectories that remain predictive under the question-only view, and Advantage-Gated Posterior-to-Prior Transfer to selectively distill validated efficient reasoning patterns into the question-only prior. Experiments across 1.5B and 7B models demonstrate consistent accuracy-efficiency improvements: VPG-EA improves the comprehensive efficiency metric by 8.5% and 17.2%, respectively, reduces token consumption by 29–44%, maintains or improves accuracy, and generalizes to non-mathematical scientific domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.