Where to Repair? World-Model-Guided Hierarchical Intervention for LLM Agents
Abstract
Effective self-correction in LLM agents requires anticipating potential failures and identifying the appropriate scope of intervention. We propose HiCARE, a framework for hierarchical counterfactual attribution and repair before execution. Our framework organizes the agent loop into hierarchical decision levels and uses a world model to simulate the outcomes of proposed actions. When a simulated outcome indicates failure, the world model evaluates counterfactual interventions at different levels to attribute the predicted failure and assess how alternative decisions could resolve it. An intervention controller balances the expected improvement against the cost of modifying each level, selecting targeted repairs within a bounded deliberation budget. Revised proposals are reassessed through world-model simulation before the final action is executed in the real environment. This process connects predictive failure attribution with counterfactual intervention, enabling anticipatory self-correction across decision levels. Experiments on DiscoveryWorld’s Combinatorial Chemistry task and -Knowledge show that our framework achieves better task performance than the evaluated baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.