DWELL: Persist or Wrap Up? Measuring and Governing Dynamical Instability in LLM Multi-Agent Orchestration
Abstract
When tools repeatedly fail, should a large language model (LLM) orchestrator persist or wrap up? Stopping early can lose recoverable successes, while continued retries can exhaust the budget on unsolvable tasks. We study this decision by separating observed progress from the allocation of further retries. We introduce a trajectory-level stability margin to measure progress and a budget-aware dwell-time controller, DWELL, that determines how long to keep retrying from the failure history and remaining budget under a simple recovery model. Experiments show that task success declines gradually as tool failures become more persistent and drops to zero when the last source of required information is lost. Comparisons with fixed stopping rules quantify how earlier stopping trades recoverable successes for resource savings. Within this trade-off, exploratory analysis on the primary frozen faulted workload finds that DWELL reduces tokens per successful task by 27.5% relative to ungoverned orchestration. Tests on new tasks with equal approach show that models stop when instructed and succeed when tools recover before the retry limit. Whether stopping helps depends on the task and how long the model would otherwise keep trying.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.