Preserve Broadly, Observe Adaptively: Budget-Decoupled Event-Point Learning for Action Recognition
Abstract
Event cameras capture motion as asynchronous events, making them well suited to action recognition under rapid motion and challenging illumination. However, event recordings vary in cardinality, while existing point-based pipelines commonly form fixed-size network inputs by randomly sampling events from each recording or temporal clip before learning. This one-shot reduction couples event retention to per-pass computation, irreversibly discarding evidence from dense recordings while padding or duplicating events in short ones. We introduce EventVista, a preservation-aware framework that decouples the retention budget from the per-pass computation budget through three advances. First, Authentic Candidate Preservation preserves authentic evidence beyond a single forward pass without increasing per-pass computation by retaining a broader, recording-wide candidate pool while maintaining the natural cardinality of short recordings. Second, Coverage-Aware Runtime Observation constructs bounded, temporally complementary views from this candidate pool and processes them while preserving their actual cardinalities. Adaptive Evidence Integration then sequentially aggregates predictions from the shared network and stops the observation process once the accumulated prediction becomes sufficiently confident and stable.Across THU-EACT-50-CHL, DVSGesture, and DailyDVS-200, EventVista achieves state-of-the-art Top-1 accuracies of 70.40% ( pp), 99.35%, and 50.68% ( pp), respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.