acceptodds
Under review as a conference paper at ICLR 2027

Where to Repair? World-Model-Guided Hierarchical Intervention for LLM Agents

Abstract

Effective self-correction in LLM agents requires anticipating potential failures and identifying the appropriate scope of intervention. We propose HiCARE, a framework for hierarchical counterfactual attribution and repair before execution. Our framework organizes the agent loop into hierarchical decision levels and uses a world model to simulate the outcomes of proposed actions. When a simulated outcome indicates failure, the world model evaluates counterfactual interventions at different levels to attribute the predicted failure and assess how alternative decisions could resolve it. An intervention controller balances the expected improvement against the cost of modifying each level, selecting targeted repairs within a bounded deliberation budget. Revised proposals are reassessed through world-model simulation before the final action is executed in the real environment. This process connects predictive failure attribution with counterfactual intervention, enabling anticipatory self-correction across decision levels. Experiments on DiscoveryWorld’s Combinatorial Chemistry task and -Knowledge show that our framework achieves better task performance than the evaluated baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.