acceptodds
Under review as a conference paper at ICLR 2027

LLM-RegProto: LLM-Enabled Cross-Modal Prototype Regularizer for Weakly Supervised Event Semantic Segmentation

Abstract

Event cameras offer microsecond temporal resolution and high dynamic range, making them attractive for autonomous perception. However, the sparsity and texture-free nature of event data render pixel-wise annotation prohibitively expensive, bottlenecking fully supervised semantic segmentation. We present LLM-RegProto, a weakly supervised framework requiring only extremely sparse point annotations (as few as 1 class 1 point, or a standard 1 class 10 points) to achieve competitive segmentation. The core innovation repurposes a frozen, pretrained text-only LLM layer as a cross-modal prototype regularizer: visual tokens from the event-stream encoder are projected into a high-dimensional linguistic manifold, where the frozen transformer layer induces semantically structured class prototypes via its pretrained attention mechanism. Unlike multimodal LLM approaches requiring paired vision-language training or text prompts, our framework feeds visual features directly into a pure-text LLM layer, with class prototypes formed autonomously through learnable class query tokens without textual supervision. To bridge the visual–linguistic modality gap, we design (i) a cross-modal adapter with 2D spatial positional encoding and bidirectional attention, enabling the autoregressive LLM backbone to process non-sequential spatial tokens; and (ii) a dual-stream prototype alignment network with background-noise suppression, jointly training two encoders on temporally complementary event windows and filtering ego-motion-induced clutter via contrastive prototype alignment. On DDD17 and DSEC benchmarks under the 1C10C budget, LLM-RegProto achieves mIoU gains of +2.16% and +2.04% over prior state-of-the-art weakly supervised methods, demonstrating that LLM pretraining weights encode universal structured priors transferable across modalities without any fine-tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.