Variational Self-Correction via Latent Diagnosis
Abstract
This paper studies post-hoc self-correction, in which an LLM revisits an initial response, identifies flaws in its reasoning, and revises the response before producing a final answer. Effective self-correction requires not only detecting what is wrong, but also translating that diagnosis into a successful revision without disrupting reasoning that is already correct. Existing methods, however, typically leave diagnosis implicit or optimize diagnosis and refinement in isolation, weakening the link between identified errors and the revisions they are intended to guide. To address this gap, we propose VaCrit, a variational self-correction framework that treats structured critique as a latent diagnosis connecting an initial response to its refinement. We first learn a critique prior through error-localization supervision and meta-utility learning, where sampled critiques are evaluated by the correction gains they induce. Building on this, we jointly optimize a critique posterior and a unified model for diagnosis and refinement under a shared variational objective, allowing refinement outcomes to assign credit to sampled critiques while preserving oracle-free diagnostic capability. Experiments on 7 challenging mathematical reasoning benchmarks show that VaCrit consistently outperforms existing methods across different model scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.