Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMs
Abstract
Closed-loop revision is increasingly adopted in large language model applications to iteratively reduce mismatches between generated outputs (e.g., a paragraph) and user targets (e.g., required paragraph length). However, revision failures may stem from incomplete feedback or models failing to act effectively on correct feedback, making it difficult to systematically assess the revision capabilities of LLMs. In this work, we propose a fixed-budget revision protocol in which deterministic verifiers report all remaining violations across three constraint families: exact-length, lexical, and compositional. This holds feedback correctness and completeness fixed, allowing us to isolate and evaluate model-side closed-loop revision. We conduct experiments on 19 mainstream open-source and closed-source models, including Llama 3.1 8B and GPT-5.6 Sol, and observe large cross-model and cross-constraint differences in final joint success, with controller-level mean final joint success ranging from 17.4% to 99.8%. Substantial cross-model gaps persist when the same initial draft is used. In controlled experiments, we find reproducible model-specific differences in how exact feedback is translated into revisions, while post-training and model scale reshape these responses without consistently bringing them closer to exact correction. Interestingly, across all constraint families, failure trajectories often repeat earlier outputs, and prior recurrence is associated with substantially lower subsequent recoverability. Finally, matched-state interventions across constraint families show that removing earlier dialogue, with the current draft and feedback fixed, changes recurrence escape without reliably improving final success; effects depend on the model, task, and trigger-state composition. Overall, exact feedback makes revision errors observable, but does not make the closed loop reliable. Code and data are available anonymously.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.