acceptodds
Under review as a conference paper at ICLR 2027

Counterfactuals, Proactivity, and Personalization: Personal HomeBench for Evaluating Smart Home Agents

Abstract

We introduce PERSONALHOMEBENCH (PHB), a benchmark for personalized smart-home agents under partial observability, comprising 9168 reactive and proactive instances grounded in multi-resident personas, long-term histories, evolving household state, heterogeneous appliances, and real-home multimodal observations. We further introduce PERSONALHOMETOOLS, a unified toolbox for context retrieval, memory access, multimodal perception, and appliance interaction. PHB evaluates models under Sole-Reasoning and Agentic settings, separating reasoning over available context from actively acquiring and integrating personalized information. Across models, we observe a pronounced reasoning-to-interaction gap, including reduced counterfactual robustness and a mismatch between plausible personalized plans and valid executable actions. Failures arise primarily from incomplete context acquisition, tool hallucination, and execution errors, while gains from scaling, explicit reasoning, and multimodal perception remain task-dependent. Overall, PHB provides a unified setting to probe how personalization, reasoning, perception, and tool use come together in realistic smart-home agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.