acceptodds
Under review as a conference paper at ICLR 2027

Redundancy-Aware Video Segment Retrieval for Few-Shot Robot Imitation Learning

Abstract

Retrieval-augmented imitation learning improves few-shot robot policy learning by augmenting limited target demonstrations with relevant segments from a large prior dataset. Existing methods, however, often rely on framewise image similarity, which can retrieve visually similar segments despite differences in the underlying manipulation behavior. Moreover, independently ranking candidates with a Top-K rule can fill the retrieval budget with redundant segments that exhibit highly similar spatiotemporal manipulation patterns, limiting the benefit of data augmentation. We propose RVSR, a redundancy-aware video segment retrieval framework that jointly considers target relevance and redundancy among selected segments. RVSR partitions robot demonstrations into temporally coherent segments using a frozen pretrained video representation and derives two role-specific features: a semantic feature for target relevance and a temporal-pattern feature for redundancy estimation. It then sequentially selects relevant segments while suppressing manipulation patterns already represented in the retrieved data. RVSR achieves 66.4% success on LIBERO-10, outperforming the strongest retrieval baseline at 63.5%, and obtains the highest average performance across three real-robot tasks. Ablations further show that temporally structured video representations and redundancy-aware selection are most effective when combined to improve few-shot policy learning. The source code is included in the supplementary materials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.