Do Embodied AI Models Have Physical Intuition? An Egocentric Probing Benchmark
Abstract
Estimating contact forces from egocentric observations requires embodied models to map visual appearance and hand motion to continuous physical quantities. These estimates depend on how object geometry, deformation, and load evolve during interaction. Occlusion and visual–physical mismatch complicate this mapping across materials and fill states. We introduce FEEL (Force Estimation from Egocentric Learning), a controlled benchmark of 3 069 vision–touch–pose episodes in ten object and fill-state categories, and EgoForce, which uses a GRU for visual–state aggregation and two softly weighted residual experts to refine a shared Laplace force head. We evaluate transformer regression policies, flow-matching vision–language–action models, generalist multimodal policies, and vision-foundation encoders on current-force prediction from RGB and hand state. The 36.18 M-parameter EgoForce achieves test contact MAE of 7.3386 N, contact , contact Pearson , and pooled F1 of 0.8395, with OOD contact and pooled F1 of 0.8934 on empty milk containers. Project page: http://submission.anonymoussubmission.com/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.