Learning Bidirectional Normative Latent Dynamics with Discrepancy Reasoning for Clinical Skill Assessment
Abstract
Vision-Language Models (VLMs) have become increasingly capable of recognizing events in videos. However, reliable procedural reasoning requires more than identifying what happened: a model must track how interactions update the underlying world state and verify whether the resulting state evolution is valid. In this work, we study this problem through clinical skill assessment, where the correctness of an action depends on the procedural state accumulated through prior interactions among the operator, patient, instruments, and environment. We propose ATV, an Anticipate-Then-Verify framework that formulates procedural assessment as verification against normative latent dynamics. ATV first organizes video observations into temporally aligned object-centric states. It then learns bidirectional dynamics exclusively from correctly performed procedures, where a forward predictor anticipates how object states should evolve and a reverse-time reflector captures whether the resulting states are compatible with their expected antecedents. At inference time, ATV converts the discrepancies between observed and normative state transitions into rubric-conditioned evidence for procedural scoring. We further construct ClinSkill-V, a large-scale dataset of 15,775 step-aligned clinical skill video clips collected from real-world clinical skill examinations, with expert rubric-based ordinal annotations. Extensive experiments against open-source VLMs and latent predictive models demonstrate that ATV consistently improves procedural assessment, with particularly pronounced gains in fine-grained ordinal scoring.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.