Measuring and Auditing Rubric-Induced Confabulation in VLM-Based Assessment of Human Procedural Skills
Abstract
Procedural skills constitute core practical capabilities with pivotal importance across K-12, vocational, and higher education systems. Nevertheless, their assessment has long relied on experts scoring students' live demonstrations against standardized rubrics, which is hardly scalable. Vision-language models (VLMs) enable scalable expert-grade observation via long-video understanding. However, we find that VLMs fail to target graded operations and frequently misjudge critical devices without rubric guidance; conversely, equipped with rubrics, the models falsely mark planned procedural steps as completed actions. We name this rubric-induced confabulation a machine analogue of rater expectation bias, and trace it to recitation: the model repeats the rubric instead of reading the clip. We therefore audit each claim: deciding whether each claimed step came from the video or the rubric, formalized as conditional pointwise mutual information (CPMI) under two counterfactual interventions (removing the rubric procedure; swapping the video), and estimated either white-box from token log-probabilities or black-box from how often each claim reappears under resampling. As an outcome-validity check, higher white-box visual grounding corresponds to lower sampled support for absent steps on all four open-weight models (Spearman , ). Computed from sampled claim recurrence alone, the trust signal needs no internal probabilities and extends to closed-source models. We release a benchmark of four domains spanning K-12 science practicals and vocational skills (874 sixty-second clips, 14k step-level occurrence labels) and evaluate nine VLMs, four open-weight and five closed-source. On the four open-weight models, the rubric raises step coverage from 22% to 68% while 60–86% of the claimed steps never happened, and the false-claim rate rises from no to full rubric for every open-weight model–domain pair (+20 points on average), with an increase in 15 of 20 closed-source model–domain cells. Rubric-recited false claims recur almost as reliably as true ones, limiting detection by cross-sample consistency; under severe imbalance where true-claim prevalence is 13–27% by domain, our trust signal achieves AP at 2.0–4.0 times the corresponding prevalence baseline on the four domain averages.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.