acceptodds
Under review as a conference paper at ICLR 2027

Choosing, Not Generating: Option Selection, Learned Preconditions, and Temporal Consistency Checks for Smart-Home Agents

Abstract

Language-model agents that operate a smart home by issuing one tool call at a time often fail on requests that involve timing, dependencies between devices, or a device or time that the user states incorrectly. Such an agent infers again at every step, from its growing history, what each command requires and when each action should happen. We observe that many of the decisions it makes are selections among options that the tool interface already lists, such as which device a phrase refers to, and that the times follow by computation from values the devices report. We therefore use text generation only to convert the request into a list of atomic goals, in one language-model call. A second model makes these selections: given a question and a list of options, it returns a probability distribution over the options instead of generating text. It maps each goal to a device, or to none when the named device does not exist, and chooses the goal's action and arguments. Code computes all times, checks with a simple temporal network whether the request's timing constraints are consistent, and either executes the request or explains the conflict. Skills learned from demonstrations, annotated with the preconditions of their actions, add the actions that others require first on SimuHome, a smart-home benchmark, and serve as the plans on TimeArena, a benchmark of concurrent time-consuming tasks. On SimuHome, the agent reaches 92.8% and 89.4% success with Gemma 4 12B and Qwen 3.5 9B, 30.5 and 31.6 points above the strongest baseline that uses the same demonstrations; with Gemma 4 12B, it makes about one backbone call per episode instead of about seven for ReAct and uses 13 times fewer backbone tokens. On TimeArena, it completes every episode that can be completed without violating the task instructions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.