HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
Abstract
Frontier coding agents solve complex tasks when given complete context but struggle when specifications are incomplete or ambiguous. The bottleneck is not raw capability, but _judgment_: knowing when to act autonomously and when to ask for help. Current benchmarks are blind to this failure mode: they supply unambiguous detailed instructions and reward solely execution correctness, so an agent that makes a lucky guess for a missing requirement will score identically to one that would have asked for certainty. We present HiL-Bench (Human-in-Loop Benchmark) to measure this _selective escalation_ skill. Each task contains human-validated blockers (missing information, ambiguous requests, contradictory information) that surface only through progressive exploration, not upfront inspection. Our core metric, Ask-F1, the harmonic mean of question precision and blocker recall, captures the tension between over-asking and silent guessing. Evaluation across SWE and text-to-SQL domains reveals a large universal judgment gap: no frontier model recovers more than a fraction of its full-information performance in settings where it needs to decide whether to ask. Failure analysis identifies key patterns: overconfident wrong beliefs with no information-gathering; inability to use human or tool feedback in problem-solving; and broad, imprecise escalation. These consistent patterns confirm poor help-seeking is a model-level, not task-specific, flaw. Fortunately, RL training on a shaped Ask-F1 reward shows judgment is trainable: a 32B model improves on both help-seeking quality and task pass rate, with gains that can transfer across SWE, SQL, and blocker types.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.