RoSci-Video: Benchmarking Robust Scientific Understanding and Reasoning on Experimental Videos
Abstract
Scientific videos from laboratory experiments often contain missing, erroneous, or conflicting data across modalities. When models fail to recognize such degradation, they risk producing confidently wrong conclusions in scientific understanding and reasoning. Existing benchmarks for large multimodal models (LMMs) do not evaluate this capability. To bridge this gap, we introduce RoSci-Video, the first benchmark evaluating LMM robustness to information degradation in laboratory videos. It covers 3000 questions across 390 videos from 13 disciplines, with perturbations categorized as Information Error or Information Loss. We define Robustness Drop (RD) as the difference between reference accuracy and rejection rate. Evaluating 32 LMMs reveals three findings: (i) all models exhibit positive RD, indicating a universal robustness gap; (ii) information loss perturbations produce higher RD than information errors, showing models struggle more with missing information; (iii) the dominant failure mode is capability-present robustness failure: models answer correctly on clean references yet commit to incorrect answers on degraded inputs, accounting for the majority of cases. Within-family robustness trends are inconsistent, indicating that scale alone does not reliably predict degradation robustness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.