acceptodds
Under review as a conference paper at ICLR 2027

Attribute-Level Data Self-Evolution for Audiovisual Instruction Tuning

Abstract

Fine-grained video understanding relies on detailed and reliable audiovisual captions, yet caption enrichment often introduces factual errors without equally fine-grained verification and correction. Even under caption-level verification, which checks each detailed caption as a whole, we observe a critical granularity mismatch, i.e., coarse feedback fails to pinpoint fine-grained errors and omissions, limiting targeted refinement. To address this mismatch, we introduce ASID-Verify, an attribute-centered framework for verifying and refining audiovisual captions. A verifier audits each caption against the video and available speech transcripts, returning an attribute-organized report that identifies detected errors and omissions and explains each issue. A refiner uses this feedback for targeted correction and completion, and the revised caption is re-audited in the next round, enabling data self-evolution without updating the annotation models. A single round of attribute-level verification already reaches an error-free rate of 52.5%, matching two rounds of caption-level verification at 52.3%, and a second round raises it to 75.4%. With this framework, we construct ASID-1M, a million-scale collection of single- and all-attribute captions for 121K videos, and train ASID-Captioner on it via supervised fine-tuning. Across seven benchmarks spanning captioning, caption-based QA, and temporal grounding, ASID-Captioner improves fine-grained caption quality over open-source baselines and rivals Gemini-3-Pro on several captioning evaluations, while also following attribute-conditioned instructions more accurately. We will release ASID-1M, ASID-Verify, the evaluation suite, and model checkpoints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.