Real-World Debiasing: From Deployment Failures to Causal Repair of Policy Shortcuts
Abstract
In practice, visual control policies can fail after deployment despite strong source performance. Prior work has identified hidden reliance on task-irrelevant source cues as one cause of such failures, even when the underlying task remains unchanged. Yet diagnosing and repairing this dependence after deployment is difficult: multiple factors may change simultaneously, and target interventions or retraining may be costly or infeasible. We call this broader failure pattern real-world bias. We further show that observational source data alone need not identify such harmful dependence, even when source performance is optimal. To address this gap, we introduce RENDER, which uses a few unpaired target failure trajectories to guide causal diagnosis and selective repair in the source simulator. RENDER reenacts candidate nuisance mechanisms to reproduce the observed failures, then tests whether task-preserving interventions change policy decisions and increase task loss. Once harmful dependence is verified, the diagnosed mechanism guides targeted adaptor training using source data while the original policy remains frozen; otherwise, RENDER abstains. The procedure requires no target interventions. Experiments across diverse visual control tasks support our analysis and show that selective repair can improve deployment performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.