acceptodds
Under review as a conference paper at ICLR 2027

AVA-Encoder: Towards Agent-Native Video Representation Learning

Abstract

Agentic video creation systems remain far from cinematic-grade quality: their underlying models struggle to plan how story, characters, visuals, and sound fit together, and high-quality creation trajectories to learn such planning abilities are scarce. Finished human films embody this knowledge, but in a tightly coupled audiovisual form that agents cannot directly read or manipulate. We propose the Agentic Video Auto-Encoder (AVA-Encoder), which works backward from finished films, turning them into agent-native representations with reconstruction itself as supervision. It encodes a film into a Film Knowledge Graph (Film KG) of structured-text nodes, typed dependency edges, and linked multimodal assets, then regenerates the film from this graph with a fixed decoder. Reconstruction residuals are distilled into textual gradients that drive dual-loop self-evolution: an outer loop learns a shared encoding policy across videos, and an inner loop refines each video's Film KG at test time. AVA-Encoder improves reconstruction fidelity by 20.7 percentage points over the strongest baseline; its learned policy outperforms a finely human-tuned policy with over 70% fewer system-prompt tokens; and Film KGs support coordinated editing and improve all tested downstream creation systems. We release the framework, a reconstruction benchmark, and a Film KG dataset.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.