acceptodds
Under review as a conference paper at ICLR 2027

MUSE-Home: Evaluating and Aligning LLM Agents for Multi-User Smart Homes

Abstract

Smart-home assistants work in shared environments, where residents have different needs, permissions, routines, and observations, while their actions affect one another through shared spaces and devices. We introduce Multi-User Smart-home Evaluation (MUSE-Home), an executable benchmark that studies four challenges created by this setting: multi-user conflict, socially conditioned proactivity, cross-resident environment forensics, and robustness against unnecessary coordination. MUSE-Home contains 6,790 training instances and a manually audited out-of-distribution test set of 721 instances. The best evaluated model achieves only 45.35% task success. We further propose Aggregating Local Interests with Global Need-weighting (MUSE-Align), a judge-based alignment framework that combines observation-bounded resident judgments with full-context weights for current needs. We use this signal to select SFT and RL tasks, and to provide a soft reward for GRPO. MUSE-Align improves Qwen3-4B from 26.63% to 31.35% under SFT and from 20.39% to 38.56% under GRPO. These results provide a benchmark and training framework for developing LLM agents that act appropriately in shared homes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.