acceptodds
Under review as a conference paper at ICLR 2027

See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making

Abstract

Reliable human-centered decision making in physical environments requires proactive agents not only to choose what action to take, but also to judge whether the current evidence is sufficient to act. This judgment is not an isolated property of the decision module; it directly reflects how deeply the agent understands the task context of the physical world. We study proactive retail service from sparse third-person video. Before a customer makes an explicit request, the agent must use limited human–object interaction evidence to choose between intervening to guide the interaction and remaining silent to continue observing. We use physical grounding narrowly: converting real human–object observations into a task-relevant retail state, not modeling low-level physical dynamics. We introduce the Proactive Intent World Model (PIWM), organized around three functional capabilities: See constructs the perceptual basis, Foresee models counterfactual consequences, and Act selects the subsequent action. Our experiments reveal that performance is poor when the agent must extract information from raw video and decide directly, but improves substantially when it receives structured inputs extracted and annotated from a professional retail perspective. Reliable intervention therefore requires role- and goal-directed attention that selects and organizes the cues that matter for service decisions. Ablations involving AIDA-stage constraints and BDI state structure further support this finding. PIWM's counterfactual prediction remains strong in standalone evaluation, yet planning methods that explicitly query these forecasts at inference time degrade sharply. Locally useful consequence prediction therefore does not yet translate reliably into better action selection. This gap may reflect incomplete process understanding, uncertainty in fine-grained single-step outcomes, and insufficient joint modeling of the physical scene and its temporal evolution. remains the hardest action in structured-state evaluation, exposing the same challenge of temporal awareness in proactive decisions. PIWM thus advances static intent recognition toward intent world modeling: an agent must organize physical observations under professional task knowledge, anticipate how candidate interventions affect interaction state, and treat intervention and non-intervention as joint decisions. Future work will introduce long-horizon interaction trajectories and temporal consequence supervision to improve sustained reasoning about evolving situations and intervention timing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.