Recover, Don't Just Continue: Learning from Feedback for Physical Failure Understanding and Correction
Abstract
Multimodal large language models are increasingly seen as the cognitive foundation for embodied intelligence, yet planning from instructions alone cannot make them reliable embodied brains. During physical execution, actions deviate from intended effects, and what can be realized is jointly constrained by the embodiment, controller, and environment. An embodied brain must therefore understand why execution deviates, recognize what the current execution system can reliably realize, and revise decisions using interaction experience. These three needs correspond to causal failure localization, capability-constrained adaptation, and experience-driven revision, making them complementary and non-substitutable. However, current embodied foundation models, despite advances in spatio-temporal reasoning, physical-world modeling, and high-level planning, largely lack these capabilities. To address this gap, we introduce Embodied-Corrector. Controlled failure–recovery data connects actions, observable consequences, failure causes, and executable recoveries, providing physical causal supervision. Evaluation-first SFT with policy-optimization internalization transfers verified evaluations into the model's own reasoning, so it assesses outcomes and failure causes before inferring recoveries a frozen executor can realize, enabling execution-capability adaptation. A CoT template requires the model to review execution history, evaluate failure causes, and plan recovery, so each execution informs the next decision without parameter updates, realizing in-context learning from interaction. This internalizes diagnostic and corrective experience into the model's reasoning. Experiments on embodied failure reasoning and planning benchmarks show stronger failure-cause analysis and corrective planning; with a frozen VLA, composite task success improves by 6.1%, and with fixed GPT-5.5 execution, embodied manipulation and navigation tasks improve by 13.9% macro-average.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.