acceptodds
Under review as a conference paper at ICLR 2027

PACT: Toward Verifiable Proactive Assistance for Wearable Language Agents

Abstract

A wearable agent that continuously observes its user must decide, without an explicit request, whether to intervene, what assistance to provide, and how to deliver it under device and timing constraints. Existing training largely relies on answer-level supervision, which provides target decisions but does not retain an executable mechanism for verifying new policy outputs. Group-relative reinforcement learning poses a second difficulty, as identical rollouts receive identical rewards and yield no advantage signal. We formalize label derivability as the property that the decision mechanism remains executable after data construction, and propose PACT, a training framework that exploits this property for proactive policy learning. PACT derives supervision from executable decision rules over measurable conditions, re-executes these rules to verify policy outputs, and forms optimization groups across label-changing counterfactual pairs to preserve reward contrast when within-input variation disappears. We further introduce POISE, a benchmark of 8559 wearable decisions across twelve scenarios, 49 tools, and nine delivery channels, grounded in sensor and contextual observations. Evaluation on POISE reveals persistent difficulty in distinguishing warranted intervention from abstention among prompting and inference-time methods. PACT improves over supervised fine-tuning and alternative RL objectives on POISE, while also yielding gains in tool selection and proactivity-score calibration on external benchmark.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.