CoHabit Arena: Evaluating Personal Preference Realization in Long-Horizon Human-Agent CoHabitation
Abstract
Completing a household task does not ensure that an embodied agent respects a user’s unstated preferences. We formalize Personal Preference Realization (PPR) as the ability to infer a user’s latent preferences from interaction, retain them across tasks, and express them in subsequent embodied decisions. To evaluate this capability, we introduce CoHabit Arena, a diagnostic evaluation environment built on Habitat 3.0. Each session preserves a synthetic user persona and interaction history across tasks without providing its ground-truth preference profile to the agent. Executed trajectories are retrospectively matched to preference-sensitive choices and associated with persona-conditioned reactions, allowing preference evidence to be traced to prior choices and their consequences. CoHabit Arena comprises 576 sessions of 15 tasks across 12 scenes and 36 synthetic personas, totaling 8,640 task instances derived from 1,080 source tasks. Its evaluation separates preference inference (mental alignment), preference-sensitive executed choices (behavioral alignment), and task success. We include E-ToM, a diagnostic baseline that maintains a persistent belief over preference axes and uses it to select subsequent plans. Experiments with standard LLM planners and a Theory-of-Mind baseline across two LLM backbones show that E-ToM-augmented planners obtain higher mental-alignment scores and lower preference-estimation error, while task-success gains are inconsistent. CoHabit Arena thus provides a controlled setting for studying cross-task preference acquisition and testing whether inferred preferences translate into embodied behavior beyond instruction following.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.