acceptodds
Under review as a conference paper at ICLR 2027

PhysHuman: Physics Aware Human and Object Interaction Video and Dynamics Joint Generation

Abstract

Human–object interaction (HOI) video generation promises scalable creation of interactive human motions for graphics and embodied applications. However, despite rapid progress, most existing models primarily optimize for RGB fidelity while overlooking physical plausibility and actionable physical signals that govern human motion, often producing interpenetrations, implausible impulses, and motions that fail to vary coherently with physical settings and interactions. To tackle this issue, we propose PhysHuman, a physics-informed approach that jointly generates HOI video and dynamics by conditioning a diffusion-based video model on text prompt and first frame. By predicting actuator torque as an auxiliary target, the model learns structured dynamic information, thereby improving the overall quality of the generated video without providing actuator torque information . Furthermore, to facilitate learning and evaluation of physics-informed HOI generation, we introduce PhysHuman-Dataset (PHD), a controllable corpus constructed in simulation by executing HOI policies while sweeping object mass, providing synchronized RGB, metrically accurate depth and actuation traces, together with geometry and physical-parameter annotations. Across various evaluations, PhysHuman improves physics consistency while simultaneously unifying the generation and perception of HOI.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.