RetryScribe: Retry-Aware Captioning for Robotic Manipulation Videos
Abstract
A successful robotic demonstration can contain missed grasps, corrective motions, and repeated attempts that a task-level caption such as “picks up the object” leaves unmentioned. Understanding these retries requires descriptions that preserve the target, timing, and local outcome of each attempt, together with its relationship to earlier attempts and the adjustments made between them. We introduce RetryScribe, a scalable framework for retry-aware captioning of robot manipulation videos. Because vision-language models alone can overlook brief or subtle retries, RetryScribe uses action logs to locate each attempt in time and video to identify the object it acts on and whether it succeeds, producing fine-grained caption supervision for 40,784 single- and dual-arm robotic demonstrations. We then distill this supervision into video-only captioners that require no action logs at inference. To evaluate retry understanding, we introduce RetryBench, a human-reviewed benchmark of 1,206 demonstrations and 10,795 questions that measures how much information about actions, attempt timing, and retry structure can be recovered from a caption. Experiments show that RetryScribe models substantially outperform their base models in capturing retries and cross-attempt relationships, on both in-domain robot-arm and out-of-domain handheld-gripper videos, while retaining general robot-video captioning ability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.