Measure, Calibrate, Refuse and Verify: An Evidence Contract for Mobile AQA
Abstract
Deployment-facing action quality assessment combines repeated observations, user adaptation, human review, physical measurements, and abstention, yet these components are often evaluated as if their evidence were independent. We formulate an *evidence contract*: every adaptive or human-assisted degree of freedom requires a matching source of evaluation independence. We instantiate the contract in a longitudinal audit of an eight-month mobile rehabilitation system and obtain four paired reversals. Same-user threshold fitting raised accuracy from to , whereas nested leave-one-user-out calibration achieved . Pooled repetition agreement of became when users received equal weight. On the same 23 videos and frozen predictions, model agreement changed from under model-visible review to against blind consensus; 51 videos were retrospectively labeled, of which 38 entered frozen-rule scoring after exclusions, yielding error-detection accuracies of and for two exercises. Finally, within-user correlation of coexisted with centimeter limits of agreement , while refusal retained user-level risk at coverage. The resulting evaluation stack—user-level aggregation, nested adaptation, blind-first labeling, measurement agreement, and coverage–risk analysis—turns apparently reliable outputs into testable stage-wise claims. It applies beyond rehabilitation wherever systems personalize, expose machine suggestions to reviewers, or withhold predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.