ReSQ: Benchmarking Questioning, Answering, and Recovery in Embodied Agents
Abstract
Embodied agents inevitably encounter uncertainty during long-horizon task execution, requiring them not only to seek missing information but also to use it effectively to recover. We introduce Recovery through Selective Questioning (ReSQ), a benchmark built from actual failures of autonomous LLM agents for studying the complete interaction loop of questioning, answering, and recovery. ReSQ-Q contains 2,160 failure-grounded question demonstrations from 678 failed ALFRED trajectories, specifying when, what, and why to ask. ReSQ-A provides 6,480 independently collected human responses and supports a fixed response oracle for scalable and reproducible interaction. ReSQ-R provides answer-conditioned recovery trajectories demonstrating how agents revise their reasoning and actions after receiving information. Across diverse LLM families and scales, question supervision generally improves task completion, with recovery supervision providing further gains. Our selective-questioning analysis reveals distinct asking regimes across models, showing how beneficial and harmful interactions jointly shape task success. We further find that interaction effectiveness depends on question timing and response source. ReSQ provides a unified testbed for learning agents that seek help selectively and recover effectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.