acceptodds
Under review as a conference paper at ICLR 2027

DRFR:Deviation Representation and Failure Risk Learning for Weakly Supervised Continual Monitoring of Robot Task Semantics

Abstract

Robots executing language-conditioned tasks may produce physically plausible actions that still violate the intended semantics. Continuously detecting such deviations and their associated failure risk is essential for reliable autonomous execution. Yet successful executions vary in strategy and timing, and local deviations can often be corrected by subsequent actions. This makes single-reference or instantaneous detection unreliable. We present a weakly supervised framework for continuous semantic failure detection that uses only task-level success/failure labels to learn process-level risk, while dynamically expanding its reference space with accumulated reliable successes. Our framework comprises three components. First, we introduce Progress-Aware Deviation Representation with Success Memory (PADR). PADR retrieves task-relevant successful trajectories using task-conditioned vision-language representations. It then aligns the current execution prefix with multiple candidate successes through open-ended monotonic alignment, producing task-relevant relative deviation signals. Second, we propose Failure Risk Learning from Execution Deviations (FRLED). By leveraging non-negative evidence injection and a learnable recovery gate, FRLED tracks the accumulation and resolution of deviation evidence and estimates failure risk during execution under task-level success/failure supervision. Finally, we introduce Continual Success Memory Expansion via Reliability and Novelty Gating (CSME). CSME accepts, buffers, or rejects candidate experiences according to task outcome, behavioral reliability, and experience novelty. Newly confirmed success patterns are then incorporated into the reference memory. Experimental results demonstrate that the proposed framework can identify semantic task failures, while further improves detection performance on subsequent tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.