BeaconAgent: Learning Capability-Aware Help-Seeking Policies for Small Language Model Agents
Abstract
Small language model (SLM) agents offer a resource-efficient alternative to large language model (LLM) agents, but their limited capabilities make them unreliable on challenging tasks. Recent work enables SLM agents to leverage stronger LLMs through self-reflection, confidence estimation, and learned routing. However, existing approaches mainly focus on assistance routing, with limited attention to the design of a fine-grained help-seeking policy. First, the critical step where assistance is most valuable may fail to be identified. Second, assistance decisions may not align with the SLM agent's capability boundary, causing unnecessary or insufficient help. Third, existing approaches often rely on delegation, increasing LLM involvement and reducing SLM autonomy. To address these challenges, we propose , which enables the SLM to execute the full agent trajectory while selectively requesting brief, targeted LLM guidance provided only as an observation. BeaconAgent learns this help-seeking behavior through two training stages. First, a teacher LLM identifies the causing trajectory failure, replaces it with an explicit ASK_HELP step, and completes the remaining trajectory with guidance. The resulting successful trajectories are used for supervised fine-tuning, teaching the SLM agent . Second, capability-aware reinforcement learning aligns help seeking with the SLM agent's capability boundary, promoting assistance when needed while reducing unnecessary help otherwise, thereby teaching the SLM agent . We evaluate BeaconAgent on math reasoning, retrieval-augmented QA, and unseen coding tasks. Results show that BeaconAgent improves task performance while reducing unnecessary LLM usage, achieving a better performance–cost trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.