Event Camera Motion–Appearance Disentanglement
Abstract
An event camera observes changes in brightness, so the signal produced by an object depends jointly on its appearance and its movement. Separating these factors could let a representation recognize the same object under different motions, or compare movements across different objects. We ask how far this separation can be obtained from existing event representations without retraining their encoders. Our starting point is simple: observations that share one factor reveal which feature changes should be ignored when representing that factor. We study this idea through paired spectral analysis and lightweight learned projections, using independently pretrained encoders for feature analysis and reconstruction-pretrained autoencoders for rendering. A controlled dataset that records every object along every trajectory allows us to evaluate the factors separately and then test their recombination. Retrieval assesses whether the codes distinguish the requested factor despite changes in the other; rendering asks whether codes from different clips can produce the corresponding event observation. The experiments show that useful factor-specific retrieval does not by itself ensure successful rendering, and that the interface between extracted codes and a renderer substantially affects the latter. Captured-event probes delimit the applicability of the controlled findings. Together, these tests distinguish information that can be extracted from a frozen representation from information that can be independently used.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.