acceptodds
Under review as a conference paper at ICLR 2027

Beyond Corrected Actions: Training LLM Agents on Textual Critiques

Abstract

Large language models (LLMs) are increasingly used for multi-turn agentic tasks that require interactions with users and environments over extended sequences of actions. A common approach to improving LLM agents is to use more capable expert models to provide training supervision, either by behavior cloning the expert's trajectories or by using the expert to correct the actions in trajectories generated by the weaker student model. However, expert LLMs can provide richer supervision than actions alone: they can directly generate textual critique that explain *why* specific actions are incorrect. We study whether training on such feedback benefits LLM agents and propose **C**ritique-Augmented **A**gentic **F**ine-**T**uning (CAFT), a framework that elicits textual expert feedback on student-generated trajectories. For each incorrect action, the expert yields both a corrected action *and* a textual critique explaining the student's mistake—which we then use as the training target for fine-tuning. Across four student models from the Nemotron 3 and Gemma 4 families on the agentic domains -Bench Airline+Retail and WebShop, adding critiques to action-level supervision improves performance by 8% on average over action-only distillation. Additionally, CAFT checkpoints can also provide a better warm start for reinforcement learning (RL) than expert-trajectory distillation. Our results demonstrate that textual critiques can strengthen supervision for multi-turn LLM agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.