BodyReady: Grounding Human-Centere Planning in Visual Preconditions and Action Effects
Abstract
Human-centered planning and assistance require more than predicting a plausible next action: they also require knowing whether the observed person can actually begin that action from the current bodily state. Existing benchmarks largely evaluate task progression, symbolic action preconditions, or agent and object feasibility, but do not isolate whether a specific person's visible bodily state permits an action or how preceding actions change what becomes possible next. Instead of treating action plausibility as sufficient, we explicitly ground action applicability in the person's hand use, posture, and target reachability, and evaluate how action-induced changes in bodily state alter subsequent readiness. We introduce , a benchmark for visual action applicability of a specific human, comprising 1,190 questions grounded in 280 exocentric images. Six complementary tasks evaluate whether an action can begin, what prevents it, which preparation can enable it, and where a short action sequence first becomes invalid. Reference answers are derived from human-reviewed visual states, explicit action prerequisites, and bounded completion effects. Across 18 proprietary and open VLMs, the best model achieves 74.96% strict overall accuracy, but only 63.85% on selecting currently applicable actions. Even when it correctly identifies an action as unready, it fails to diagnose the correct set of bodily blocker categories in 28.88% of those cases. Models also choose preparations that cannot themselves begin and fail to identify the first invalid step when the same actions are reordered. These results reveal a gap between predicting plausible actions and reasoning about whether a particular person can execute them from the current bodily state, establishing visually grounded human action applicability as an important capability for planning and assistance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.