acceptodds
Under review as a conference paper at ICLR 2027

TWILL: A Proactive Agent That Decides from Internal Proposals and Human Requests

Abstract

A household robot should recognize from its own observations what a situation requires. Instruction-driven manipulation systems and proactive assistants that anticipate human needs share a blind spot: the robot seldom forms its own judgment of the scene, let alone carries it through to execution. A robot with its own judgment must also arbitrate between that judgment and what the person asks. We build TWILL, a proactive robotic agent that closes the loop from perception to action. Its proposer forms a proposal from the scene alone. Its decision module, SHUTTLE, arbitrates between the proposal and the human request inside a pretrained vision-language model. SHUTTLE reads the scene and each intent through separate attention paths, so that the decision distinguishes the robot's intent from the person's. An off-the-shelf executor then carries out the decision. To evaluate the full loop, we introduce TWINE, which scores understanding and decisions on 360 cases in 180 controlled pairs, each differing in one consequential factor, and scores closed-loop success in simulated kitchens. The proposer reaches 89.88% accuracy in deciding whether and what to propose, 48.61 points above the untuned backbone. SHUTTLE raises the rate of fully correct decisions by 3.70 points over the same fine-tuned backbone without separate reads. In closed loop, TWILL reaches an 81.48% success rate on request-free episodes, against 50.00% for an untuned vision-language model commanding the same executor. TWILL thereby allows a household robot to act on proposals it forms from the scene and to arbitrate between these proposals and human requests according to context.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.