acceptodds
Under review as a conference paper at ICLR 2027

Don't Learn When to Look: Post-hoc Counterfactual Gating for Visuomotor RL

Abstract

A visuomotor policy that also senses its own body pays for the camera at every step in our architecture, 98.8% of policy computation, whether or not vision is useful right now. The natural fix is a learned gate that decides when to look. We show that every in-loop version of this idea fails, and that a simple change works immediately: don’t learn when to look; infer it from the frozen policy. The diagnosis comes first. The gradient-corruption premise behind gated fusion is unsupported. Teacher-relative value gaps measure how well the student learned, not whether vision helps, and no threshold repairs them. A corrected counterfactual label produces striking domain separation which mixed-domain training, a protocol we propose, exposes as gate-critic coevolution. Three further mechanisms (soft attention, expert routing, a reward-level observation cost) either cannot discriminate, or weight vision inversely to its utility, a cold-start lock-in that changing visual difficulty does not move. Then the remedy. Train one any-gate policy with mask dropout; gate it post hoc with counterfactual signals read from the frozen network. A performance-difference bound attaches to the action counterfactual and yields a frontier with guaranteed loss. Empirically, across three seeds and four domains, post-hoc gating sits above the always-on/quota convex hull, beats a random-episode control by 12–27pp at matched savings where vision is essential, is return-neutral (98%) at 17% savings, and in mixed domains outperforms even a domain-label reference; a distilled proprioception-only predictor gates without looking at 39% real savings and no loss. No labels, no retraining. When to look is not something a policy should learn; it is something to infer from one after training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.