Event-Guided Micro-motion Audio Generation
Abstract
Traditional Video-to-Audio (V2A) generation typically suffers from temporal aliasing, as standard frame videos are “blind” to the high-frequency micro-motion essential for physically aligned acoustic synthesis. To bridge this gap, we propose MicroFoley, a generative framework re-encoding inter-frame motion into high-frequency event streams. Specifically, by explicitly disentangling static appearance cues from subtle event-driven micro-dynamics, and by grounding temporal activation in physically plausible interaction patterns aligned with human auditory perception, our method successfully eliminates acoustic hallucinations. We introduce a Bidirectional Mamba architecture to model continuous acoustic evolution from causal onsets to acausal decays, which subsequently introduces an Optimal Transport-Conditional Flow Matching (OT-CFM) decoder optimized via a designable Peak-Aware Contrastive Alignment objective. Furthermore, we construct FoleyBench, a novel audio-video-event dataset (160h) integrating simulation and real-world event streams including 450 categories with our processing pipeline. Extensive results demonstrate that our method significantly outperforms state-of-the-art baselines in generation fidelity and synchronization, achieving optimized inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.