Entropy Flow Graphs: Analytic Spatio-Temporal Graph Representations for No-Reference Video Quality Assessment
Abstract
Video quality failures have spatio-temporal structure, yet no-reference video quality assessment (NR-VQA) collapses every failure into a single scalar. We introduce the Entropy Flow Graph (EFG), a quality representation that fixes its structure analytically and confines learning to three matrices. Region nodes carry a five-cue entropy decomposition, spatial edges measure the conditional entropy deviation of a region’s co-statistics from natural regularities, and motion-compensated temporal edges let the aligned past of the video serve as its own reference, so the representation is simultaneously region-first, temporal, no-reference, and localizing, with edge weights in closed form. The analytic structure yields frame-rate invariance, permutation invariance, a dominance decomposition separating global from transient degradations, and plug-in consistency with an explicit bin-count rule. On the region-labeled D-CQA corpus EFG attains a global SRCC of 0.897 against 0.831 for the strongest holistic baseline with a localization AUC of 0.823, transfers zero-shot to KoNViD-1k and LIVE-VQC, and, serialized as structured context, lifts the region-level quality reasoning of a multimodal LLM by 12.6 points, more than a learned categorical graph under the identical protocol. We release splits, metrics, and code that make D-CQA a reusable benchmark for spatio-temporal quality localization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.