Probing Material Cues in Video Foundation Models
Abstract
General-purpose video foundation models (VidFMs) generate realistic videos that capture how surface appearance evolves across viewpoints and illumination conditions, suggesting that their internal representations may encode intrinsic material information even without explicit material supervision. Motivated by this observation, we investigate the material information embedded in pretrained VidFMs through a lightweight probing framework. We keep the foundation models frozen and measure how well material properties can be recovered from their intermediate representations. Through systematic experiments across synthetic and real-world scenes, together with feature-source ablations, we find that VidFM features provide material cues beyond those available from strong image representations. In controlled settings requiring fine-grained material discrimination, video features show a pronounced advantage over DINOv2 features, while the two representations provide complementary information in more complex scenes. Further analyses show that material recoverability benefits from broader viewpoint coverage and is influenced by object-level semantics. Qualitative synthetic-to-real results further suggest that these material cues remain useful beyond the supervised training domain. Together, these findings provide evidence that large-scale video generative models learn reusable representations containing material-relevant information and characterize the conditions under which these representations offer advantages over image features.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.