When Agents Refuse to Ask for Help: Privileged Self-Distillation for Help-Seeking Software Agents
Abstract
Language-model agents increasingly execute long-horizon software-engineering tasks, yet some tasks require information present neither in the issue text nor anywhere in the repository. Exposing an ask_human tool does not by itself induce reliable help-seeking. Using HiL-Bench, we first localize the failure by inserting two diagnostic setups between \askhuman and \fullinfo that withhold resolutions and supply only a cue about what is missing. \askdescription recovers and of the \askhuman to \fullinfo Pass@3 gap for GLM-5.2 and Qwen3.8-27B, while \askhint recovers and . Identifying which information is missing, rather than acting on it once supplied, is therefore the dominant measured bottleneck. We then ask whether cue-elicited behaviour transfers to a model that receives no cue at inference time. We construct \dataset, a corpus of blocker-augmented tasks built from SWE-ReBench V2 by rewriting only the observable task specification while leaving the repository, revision, test suite and target patch unchanged. We introduce Privileged Self-Distillation (\psd), in which a teacher rolls out with privileged blocker cues, the cues are removed, and an unprivileged student initialized from the same weights is trained on the resulting trajectories. We study a matrix of privileged cue source (\askhint versus \askdescription) and supervision policy (\askall versus \askclean, which masks the training loss on judge-rejected asking turns while leaving those turns in context). All four recipes fine-tune Qwen3.8-27B with reasoning enabled. \descclean attains the highest task success (Pass@1 , Pass@3 ) and the highest blocker recall (), whereas \descall attains the highest question precision () and Ask-F1 (). Cleaning raises task success and recall under both cue sources, and lowers precision and Ask-F1 under both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.