PhysProbe: Probing Physical-Map Recovery in Frozen Vision Generative Models
Abstract
Physical perception requires recovering the scene factors behind an image, beyond producing a plausible RGB transformation. Can frozen general-purpose image editors expose these factors through their native interfaces? We introduce PhysProbe, a protocol-aware framework that evaluates depth, surface normals, albedo, roughness, and metallicity against dense physical references. Each score describes a model-protocol pair with fixed inputs, output conventions, metrics, and aggregation. The results reveal three separations: local depth-boundary recovery does not imply accurate relative geometry; low full-map material error does not imply successful localization of sparse metallic regions; and additional context does not uniformly improve physical recovery. Specialists retain large advantages in relative geometry, surface orientation, and albedo. Gemini achieves the strongest reported roughness metrics under the evaluated protocol, while metallic evaluation reveals that full-map fidelity and sparse-region localization can favor different models. Normal exemplars and auxiliary cues can further change observed rankings across models, showing that physical recovery depends on both target properties and access protocols rather than a single model ordering. PhysProbe probes the physical-perception component of Physical AI while separating observed output fidelity from hypotheses about internal representations or training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.