acceptodds
Under review as a conference paper at ICLR 2027

Event-Guided Micro-motion Audio Generation

Abstract

Traditional Video-to-Audio (V2A) generation typically suffers from temporal aliasing, as standard frame videos are “blind” to the high-frequency micro-motion essential for physically aligned acoustic synthesis. To bridge this gap, we propose MicroFoley, a generative framework re-encoding inter-frame motion into high-frequency event streams. Specifically, by explicitly disentangling static appearance cues from subtle event-driven micro-dynamics, and by grounding temporal activation in physically plausible interaction patterns aligned with human auditory perception, our method successfully eliminates acoustic hallucinations. We introduce a Bidirectional Mamba architecture to model continuous acoustic evolution from causal onsets to acausal decays, which subsequently introduces an Optimal Transport-Conditional Flow Matching (OT-CFM) decoder optimized via a designable Peak-Aware Contrastive Alignment objective. Furthermore, we construct FoleyBench, a novel audio-video-event dataset (160h) integrating simulation and real-world event streams including 450 categories with our processing pipeline. Extensive results demonstrate that our method significantly outperforms state-of-the-art baselines in generation fidelity and synchronization, achieving optimized inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.