Matched Counterfactual Attribution: A Controlled Evaluation of Data Selection for Video Instruction Tuning
Abstract
Selecting instruction data is an increasingly important way to control multimodal post-training cost, yet reported gains are difficult to interpret when subset size, baseline choice, training variation, and evaluator behavior change together. We revisit this controlled policy-comparison setting with selector-specific matched counterfactuals, a frozen 10,656-example budget, and a crossed subset-draw by training-seed design. The design separates the total policy contrast (S-U), the shift into the realized data and processing envelope (M-U), and residual ranking within that envelope (S-M). We apply it to a frozen mixed image-video pool with Qwen3-VL-8B training and response-invariant evaluation on six benchmarks. In a homogeneous-H20 3-draw by 3-seed study, ScalSelect yields +1.137 percentage points (pp) for S-U (95% CI [+0.312, +1.963]), but its S-M estimate changes from -0.233 pp under three arbitrary diagonal pairs to +0.263 pp under the crossed design (CI [-0.131, +0.657]). Thus the pairing reverses the attribution estimate without establishing a positive ranking effect. A mechanism-distinct, homogeneous-A800 XMAS adaptation gives +0.530 pp for S-U and -0.009 pp for S-M, with wide intervals spanning both signs. ScalSelect intervals use three training-seed marginals after averaging three control draws within seed; XMAS intervals use 5 marginals after averaging 5 draws within seed. The results do not rank or reject selectors in general. They show that a selector gain is interpretable only relative to an explicit counterfactual resolution, stochastic design, and training-evaluation contract.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.