OneDay: Towards Full-Day Personalized Embodied Agents
Abstract
Embodied agents powered by multimodal large language models (MLLMs) have shown strong open-vocabulary understanding and long-horizon planning. Yet, existing benchmarks largely focus on single-episode instruction following, where goals are fully specified and horizons are short. In contrast, practical household assistance demands stateful, day-level support that respects user preferences, satisfies temporal constraints, and adapts to observations shaped by earlier activities. To close this gap, we introduce OneDay-Bench, a benchmark for full-day personalized assistance, consisting of 1,000 diverse daily plans paired with structured user profiles and human-annotated preference memories. Our benchmark provides comprehensive metrics spanning personalization alignment, task completion, and time management, and validates action executability via grounded interaction in simulation. Building on this benchmark, we present OneDay-Agent, an MLLM-driven embodied assistant that executes underspecified natural-language daily plans through checkpoint-level interaction in evolving home environments. The agent performs schedule-aware goal synthesis by jointly reasoning over the plan, current observations, and retrieved user preferences, producing temporally ordered and executable actions. Together, we aim to establish a new baseline for genuinely personalized, day-long embodied assistance. All materials will be open-sourced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.