CIRRA: Dual-Level Continual Instruction Reconciliation with Ongoing Execution for Embodied Robot Agents in Interactive Household Tasks
Abstract
As the demand for assistance with household chores continues to grow, household robots have become a major focus of embodied intelligence. Real-world household deployment, however, inherently requires continuously replanning on human–robot interaction, as users may issue new instructions while previous requests are still under execution. Existing robot agents commonly accommodate such instructions by regenerating or extensively revising the remaining task sequence, which can introduce plan ambiguity, logical inconsistency, and redundant execution. To address these limitations, we formulate continual instruction reconciliation and propose CIRRA (Continual Instruction Reconciliation for Robot Agents), a dual-level framework combining LLM-based semantic reasoning with rule-constrained structural integration. CIRRA first uses semantic reasoning to determine whether an incoming instruction can be grounded to a unique executable skill and whether its required execution location is specified, resolving underspecified action and location information when necessary. It then preserves the ongoing subtask sequence as an execution backbone and generates fine-grained integration candidates by inserting the incoming subtasks into location-matched segments of the current task chain. Finally, the semantic reasoner evaluates only the modified segments to identify task dependencies and potential conflicts and select the most logically coherent local integration. By coupling structure-preserving local integration with semantic grounding and logical selection, CIRRA maintains alignment with the robot’s current execution, mitigates plan ambiguity and logical inconsistency, and reuses shared subtasks to reduce redundant execution. We further introduce CHIRP (Continual Household Instruction Reconciliation and Planning), a text-based benchmark of 120 episodes spanning eight household environments and six categories of everyday activities. On CHIRP, CIRRA reaches 74.2% decision agreement, 30 points above the strongest replanning baseline, and every correct fusion decision it makes yields a correctly placed, conflict-free schedule. On a Unitree G1 humanoid, CIRRA interrupts ongoing skills at the correct moment in every trial and significantly outperforms all baselines on every metric.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.