Evaluating Semantic State Response and Generated Visual Feedback under Distinct Input Contracts
Abstract
Interactive video systems must do more than predict action-conditioned states: their generated observations may become inputs to later decisions. We introduce State–Feedback Separation (SFS), a contract-based reporting schema for factual semantic response (S), no-anchor generated feedback (G), privileged-reference dependence (A), and transfer stress (T), illustrated by a bounded case study and mapping diagnostic. In inspected, scene-disjoint AI2-THOR FloorPlan5 evaluations, a post-hoc ResNet-50 probe reaches source/target accuracy 0.750/1.000, but an action-only rule reproduces all 160 target labels. Two distinct no-anchor RGB routes fail their gates: an object-aware predictor scores 0.375–0.500 with nonpositive gain, and a ResNet-50-guided full-resolution editor scores 0.375–0.500 across five seeds, also with nonpositive gain. Both are weak with factual histories, so these results do not isolate damage from feedback. In a separate anchor diagnostic, a distinct editor scores 200/200 with its factual anchor versus 100/200 with a blank anchor; one anchor-dropout variant preserves the pattern (0/5 seeds meet criterion). Transfer gates fail on Toaster (0/20), StoveKnob (1/20), and Kettle (1/20); oracle-goal inputs raise all tested cells to 1.000, localizing failure to goal inference in this controller pipeline. Both EPIC-KITCHENS routes are negative (0/5). On a separate controlled RGB task, post-hoc calibration yields 0.719 success (gain 0.250 over factual copy) across five frozen decoders, but passes only 1/5 fresh-training seeds. S and G use different systems. The results do not establish a within-checkpoint dissociation, internal causal mechanism, broad cross-architecture or natural-video generality, or a complete interactive world model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.