JACE: Judge with Attribution, Calibration, and Evaluation for Multimodal Audio-Video Generation
Abstract
Video generation increasingly combines multimodal conditions to produce visual and auditory content, placing greater demands on evaluation. Existing evaluators struggle to jointly support diverse generation tasks and provide interpretable judgments aligned with human preferences. Accurate defect prediction requires understanding physical plausibility, temporal coherence, and cross-modal consistency, which is challenging. Even accurate diagnosis, however, does not determine preference: shared defects may offer little basis for comparison, while different defects may carry different importance. We introduce JACE (Judge with Attribution, Calibration, and Evaluation), a generalist evaluator that connects defect diagnosis to human preference through explicit scoring contributions. JACE first learns defect predictions from non-exhaustive annotations, using complementary representations for category-specific refinement. It then freezes the diagnostic module and learns defect-category weights from separately collected pairwise preferences, without requiring both annotation types on the same videos. Combining weighted defect contributions with a base quality score yields interpretable win, tie, and loss judgments. JACE supports text, image, audio, and video conditioning across text-to-video, image-to-video, reference-to-video, and video editing tasks, jointly assessing generated visuals and sound. Its shared scoring mechanism connects individual-video diagnosis and scoring, candidate comparison, and generator ranking, supporting evaluation across generators. Across four generation tasks, JACE significantly improves defect Recall@3 (+24.6%) and non-tie accuracy by (+8.09%) over the strongest baseline (Gemini 3.7 Flash).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.