Robustifying Decision Transformers at Test Time with Direct Compensation
Abstract
Policies trained offline, such as the Decision Transformer (DT), degrade when the deployment transition dynamics differ from those seen during training. adaptive control offers a test-time remedy, but its cancellation step is built for control-affine dynamics, so applying it to a learned model has meant affinizing that model and tuning when to refresh the affinization. We propose direct compensation, an -inspired compensator that uses the learned next-state model directly with no affinization or tolerance needed. Analytically, direct cancellation's undamped first step is a regularized affine inversion, and when the learned model is control-affine and no regularizer is used, the direct cancellation coincides with the affine inversion. On MuJoCo locomotion under matched disturbances, direct overall achieves better compensation than affine for DT policies trained across diverse environments, and datasets. Ablation studies locate the primary source of the gain in the removal of the stale affinization action, which is made possible by direct 's non-interrupted state predictor and filter updates. The approach is agnostic to the base policy, so it applies beyond transformer-based policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.