acceptodds
Under review as a conference paper at ICLR 2027

TGPED: Task-Generalized Procedural Error Detection and Explanation using Large Language Models with Procedural Grounding

Abstract

In this paper, we address temporal action segmentation, error detection, and error explanation within a single unified framework across multiple procedural tasks. Existing approaches typically require a separately trained model to segment actions for each task, limiting their scalability and practicality. Moreover, due to the limited availability of training videos containing errors, prior methods are often trained only on normal videos, preventing direct optimization for error detection and explanation. To address these limitations, we propose TGPED, a task-generalized framework for procedural error detection and explanation that avoids task-specific retraining while enabling direct optimization for error understanding. First, TGPED generates dense video captions and introduces Procedure-Grounded Temporal Segmentation (PGTS), which performs segmentation using Large Language Models (LLMs) together with generated procedural references that encode task-specific information, such as action semantics and valid action orderings. Second, we design a data generation pipeline to synthesize realistic sequences of erroneous actions, forming a cross-task training dataset. Finally, we formulate the generated sequences as inputs for fine-tuning an LLM-based error detector to jointly perform error detection and explanation. We evaluate TGPED on EgoPER and CaptainCook4D, achieving state-of-the-art performance on EgoPER and strong results on CaptainCook4D. Code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.