Resource-Aware Reasoning Orchestration for Decision-Making in Embodied Agents
Abstract
Embodied agents built on multimodal large language models (MLLMs) increasingly rely on explicit reasoning, such as replanning and verification, to recover from execution failures and resolve ambiguity. However, invoking such reasoning uniformly is wasteful: many decisions in a household task are unambiguous and do not benefit from additional deliberation, while others critically require it. Existing systems address this trade-off with fixed strategies, always reasoning, never reasoning, or reasoning on a static schedule, none of which adapt to the evolving execution state. We introduce RARRL (Resource-Aware Reasoning via Reinforcement Learning), an orchestration framework that learns when to invoke reasoning rather than prescribing it a priori. A lightweight policy, trained once with reinforcement learning in a Abstract Decision Environment (ADE), observes a compact state summarizing task progress, execution reliability, and remaining budget, and decides at each timestep whether to act directly or invoke a Planner-Verifier reasoning procedure before acting. Deploying this policy on a new simulated platform requires an environment-specific grounding adapter that exposes this shared state; the policy transfers zero-shot without retraining. We evaluate RARRL on EB-ALFRED and EB-Habitat, two embodied benchmarks built on distinct simulators with non-overlapping action spaces, and find that it consistently outperforms never-reasoning, always-reasoning, and fixed-schedule reasoning baselines across six task categories, while invoking reasoning substantially less often than always-reasoning agents.We further evaluate RARRL on VirtualHome and ALFWorld, achieving the best task success rate. We further perform matched-state counterfactual interventions in the target environments, showing that the value of invoking reasoning varies across execution-time decision states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.