The Right Advice at the Right Time: Learning Proactive Assistance through Nested Credit Assignment
Abstract
In complex interactive tasks, users may spend considerable time pursuing ineffective strategies before pausing to seek external help. Proactive assistance can address this gap by offering contextual guidance without waiting for explicit requests or taking over task execution. Learning such assistance involves two coupled decisions: when to intervene and what advice to provide. Whether intervention is worthwhile depends on the advice provided, while evaluating that advice must account for the user’s ability to progress without assistance. Downstream task outcomes therefore provide ambiguous feedback for learning intervention timing and advice content. In this work, we propose PACE Proactive Assistance through nested Credit Estimation, a reinforcement learning framework that jointly trains a generative vision–language companion through nested credit assignment. By comparing estimated returns under silence and alternative advice from the same interaction state, PACE derives distinct learning signals for the intervention decision and conditional advice generation. These signals guide joint optimization of timing and content while balancing task progress against intervention cost. Experiments with multiple LLM-controlled players in ALFWorld and Minecraft show that PACE achieves higher overall task success rates with lower intervention rates than general-purpose LLMs and specialized game-playing models used as companions. A human study further demonstrates the effectiveness of PACE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.