What Should a Streaming Coach Remember? Motion Evidence Across Repetitions and Transitions
Abstract
A streaming coach is a vision-language model that watches workout video as it arrives and responds proactively. After it asks for a deeper squat, judging the next repetition requires the depth that prompted the correction; once the person moves on to lunges, that squat-specific guidance must no longer apply. It therefore needs movement evidence that stays available across repetitions and whose relevance changes at exercise transitions. Existing coaching models are trained on short single-exercise segments, and processing a longer stream does not itself keep earlier movement available for later feedback. We ask what a streaming coach should remember and propose StreamingCoach, which represents movement in the terms a coach reasons in and retains it independently of the feedback it generates: Form Tokens encode each chunk in anatomical joint coordinates grounded in person-centered visual features, and a gated Motion State is updated from them at every chunk, whether or not feedback is generated. On QEVD workouts processed as continuous streams without context reset at exercise boundaries, segment-trained coaches break down beyond a single exercise, and StreamingCoach improves reference-based feedback content and response timing over a streaming backbone fine-tuned under the same protocol, with both components contributing; on separately adapted probes it also recognizes the current exercise and counts repetitions substantially better. These results expose coaching demands that exercise-segment evaluation does not capture and turn what a streaming coach should remember into a testable question for streaming vision-language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.