acceptodds
Under review as a conference paper at ICLR 2027

RIVET: Multi-Evidence Relational Inference for Rare Surgical Triplets in Sinus Endoscopy

Abstract

Recognizing surgical interactions as instrument–verb–target (IVT) triplets is hindered by long-tailed class distributions and visual ambiguity among triplets that share components. These challenges are particularly pronounced in sinus endoscopy, where narrow operative fields, occlusions, and visually similar tissues obscure discriminative local cues. Here we introduce SinusTriplet-95, a dual-center dataset of 324 videos from 15 patients, comprising 670,381 frames and 2,428 segments annotated by multiple senior attending surgeons across 95 valid IVT triplets. We further propose RIVET (Relational Inference with Visual Evidence for Tail Triplets), which refines long-tailed triplet prediction at inference by integrating structural, historical, and local evidence. Its Dynamic Structural Relation module bidirectionally calibrates component- and triplet-level predictions. To improve the recognition of underrepresented triplets, the Tail-Class Memory Bank stores observations of rare triplet classes across surgical process, while the Tail-Class Prototype Query searches the current frame for local visual evidence associated with each rare triplet candidate. These complementary sources of evidence are integrated through frame-adaptive gating. On SinusTriplet-95, the RIVET ensemble achieves an of 50.80%, surpassing the strongest ensemble baseline by 5.71 percentage points and ranking first among the compared methods across all six AP metrics. Progressive ablations show incremental gains from each evidence source, while class-level analysis shows improvements over the corresponding baselines in 46 of 50 tail classes. Together, SinusTriplet-95 and RIVET provide a benchmark and an evidence-guided framework for surgical interaction recognition in sinus endoscopy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.