Benchmarking Egocentric Cooking Video Understanding
Abstract
Recent advances in multimodal large language models (MLLMs) have pushed video understanding toward real-world assistance. Egocentric cooking is a demanding testbed for this goal, requiring fine-grained hand-object recognition, evolving state tracking, and sequential reasoning, yet existing video QA benchmarks largely emphasize general-purpose understanding or assume access to complete videos rather than past-only observations. To fill this gap, we present EgoChef, a benchmark comprising 61,328 open-ended question-answer pairs from 129 real-world videos and covering current-state perception, temporal reasoning, and operation-level understanding. Each question is anchored to a timestamp, with models restricted to video observed up to that point and future frames withheld. For reliable open-ended evaluation, we provide multiple reference answers through multi-model semantic fusion with human calibration. Across representative open- and closed-source MLLMs, the strongest model achieves only 55.68% overall accuracy, with temporal reasoning remaining particularly challenging. We believe EgoChef can serve as a focused benchmark for fine-grained and temporal understanding of egocentric cooking.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.