acceptodds
Under review as a conference paper at ICLR 2027

Learning Bidirectional Normative Latent Dynamics with Discrepancy Reasoning for Clinical Skill Assessment

Abstract

Vision-Language Models (VLMs) have become increasingly capable of recognizing events in videos. However, reliable procedural reasoning requires more than identifying what happened: a model must track how interactions update the underlying world state and verify whether the resulting state evolution is valid. In this work, we study this problem through clinical skill assessment, where the correctness of an action depends on the procedural state accumulated through prior interactions among the operator, patient, instruments, and environment. We propose ATV, an Anticipate-Then-Verify framework that formulates procedural assessment as verification against normative latent dynamics. ATV first organizes video observations into temporally aligned object-centric states. It then learns bidirectional dynamics exclusively from correctly performed procedures, where a forward predictor anticipates how object states should evolve and a reverse-time reflector captures whether the resulting states are compatible with their expected antecedents. At inference time, ATV converts the discrepancies between observed and normative state transitions into rubric-conditioned evidence for procedural scoring. We further construct ClinSkill-V, a large-scale dataset of 15,775 step-aligned clinical skill video clips collected from real-world clinical skill examinations, with expert rubric-based ordinal annotations. Extensive experiments against open-source VLMs and latent predictive models demonstrate that ATV consistently improves procedural assessment, with particularly pronounced gains in fine-grained ordinal scoring.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.