acceptodds
Under review as a conference paper at ICLR 2027

Towards a Systematic Study of Embodied System 2 over Full Task Execution Cycles

Abstract

An embodied brain, or System 2, must turn an abstract goal into actions while discovering relevant scene conditions and checking whether executed skills achieved their intended effects. Open-loop evaluations cannot test observations shaped by an agent’s decisions, while some closed-loop evaluations provide limited evidence during execution or assume idealized skill outcomes. We introduce SCOPER, a suite for studying observation-grounded planning, execution, and recovery. SCOPER-bench combines high-level goals, ongoing visual observations, a shared rule-based System 1 for low-level actions, and controlled execution failures. Its construction and validation pipeline yields 870 tasks across BEHAVIOR and ProcTHOR, complemented by about 3,500 situated questions tied to the evaluated agent’s execution. Benchmark results reveal difficulties in retaining scene evidence, monitoring progress, and recovering from unexpected outcomes. To address these difficulties, SCOPER-harness combines Evidence Memory and Progress Verification with a dynamic task agenda and bounded exploration, without additional model training. Across two evaluated open-source backbones and two simulators, the harness improves task success by 11.1–28.6%; a harnessed 9B model outperforms a directly executed 27B model on both simulators. In an offline next-action evaluation using observations from 200 real-robot demonstrations, the harness also improves full-call accuracy by 14.7% and 15.3%. SCOPER provides a setting for diagnosing and improving embodied reasoning throughout a complete task execution cycle.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.