acceptodds
Under review as a conference paper at ICLR 2027

DrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis

Abstract

Current automated gait analysis for clinical applications often relies on black-box classifiers with limited interpretability. Although Vision-Language Models (VLMs) offer strong reasoning capabilities, applying them directly to gait videos can produce unsupported explanations, because they struggle to measure subtle geometric deviations from raw visual contexts. To address this, we introduce DrGait, a training-free agentic framework that shifts the VLM's role from a direct visual reasoner to a clinical planner. DrGait decouples semantic reasoning from geometric perception through a structured Triage-Verification-Synthesis (TVS) workflow. Given an input video and a set of basic spatiotemporal metrics, the DrGait agent first performs a heuristic triage to propose hypotheses on gait patterns, which are then verified by autonomously calling deterministic biomechanical tools that operate on reconstructed 3D mesh trajectories, segmented 2D pose tracks, and event-centered video evidence. Finally, a closed-loop mechanism recursively updates the agent's reasoning context based on the feedback. By anchoring VLM reasoning in verifiable geometric and temporal measurements, DrGait reduces hallucinations and improves gait classification, while generating transparent and interpretable clinical reports that greatly assist in clinician's diagnosis in practice.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.