Regression in Repair: Understanding Plan Repair Dynamics in Large Language Models
Abstract
Recent advances have shown that adding feedback and tools enables large language model (LLM) agents to solve increasingly complex planning problems across diverse domains. One effective paradigm for plan repair involves LLM planners producing a candidate plan, with programs checking every constraint of this plan and feedback or tools fixing the constraints it violates. In this paper, we study plan repair under coupled constraints, which price every choice against one budget so that every infeasible option of a slot is cheaper than every feasible one. We systematically analyze repair dynamics across three dimensions, coupling, the content of the repair message, and the order of the options, through empirical studies on 4,800 synthetic tasks, a second environment with requirements in text, and 900 TravelPlanner plans using two open-source models (4B and 8B parameters), gpt-4.1 and gpt-4.1-mini. Our experiments reveal three key findings about repair effectiveness when models answer in one step, namely (1) Uncoupled tasks allow feedback to more reliably fix failed attempts without breaking another constraint; (2) Exact totals from a calculator tool produce repairs that break fewer constraints than messages that only name the failed check; (3) Option choice is generally correlated with price, but this relationship varies with coupling, under which the cheaper choice is the infeasible one. These findings reveal opportunities for optimizing basic repair strategies in planning applications. First, given the same model, a single call on options sorted by the shared resources can nearly match a calculator tool in pass rate (e.g., with price and duration sorted by their weighted sum, the gap to the calculator tool shrinks from 20.0 and 15.0 points to 1.0 and 4.5 points at a quarter of its cost). Second, we identify cases where four samples at temperature 1.0 offer limited advantages over one attempt, as all samples read the same option order, suggesting that sampling alone cannot overcome position bias.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.