ActiveBrain: Benchmarking and Learning Active Cognition in Embodied Agents
Abstract
Can embodied agents actively determine what remains unresolved throughout an interaction and act accordingly to complete the original instruction? This requires seeking information beyond the current view, distinguishing the instruction-specified target from similar instances, and revising ineffective behavior. We call this Active Cognition (AC): actively updating decisions based on new observations and action consequences while preserving the instructed goal. To evaluate AC, we introduce ActiveBrain-Bench, comprising 1,035 tasks across six tracks with a common policy interface. Its Adapt track separately measures local progress and instruction completion. Zero-shot evaluations of 18 open-weight models and GPT-5.5 reveal fragmented strengths across the six tracks. In Adapt, models may make local progress yet leave the instruction unfinished. To study whether this capability can be learned, ActiveBrain constructs 11.5K replay-verified continuations from generated start states and rollout failure states. Each continuation retains RGB–action history and ends in simulator-verified completion. Our post-training pipeline combines supervised fine-tuning, on-policy distillation, and multi-turn reinforcement learning with environment-derived rewards. ActiveBrain-4B raises Qwen3.5-4B's unweighted six-track mean success rate from 8.1% to 44.0%. ActiveBrain-2B reaches 35.0%, above all evaluated zero-shot models. Across seven general multimodal benchmarks, ActiveBrain-4B records an unweighted mean of 72.6, versus 71.1 for its Qwen3.5-4B initialization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.